What you will learn
- What EBS is and how it differs from local disk and network file systems
- All six EBS volume types — gp3, gp2, io1, io2, st1, sc1 — and the single most important gp3 vs gp2 distinction
- EBS Multi-Attach for io1 and io2 — the narrow use case and its constraints
- The four-step process to encrypt an existing unencrypted EBS volume
- EBS Snapshots — incremental backups, cross-region copy, and how to share across accounts
- Recycle Bin — recovering accidentally deleted snapshots before they are gone forever
- Fast Snapshot Restore — pre-warming snapshots to eliminate initialisation latency
- EC2 Instance Store — ephemeral, highest IOPS, and exactly when to use it over EBS
- The fundamental throughput and IOPS differences that drive every volume type decision
Why this matters
Zerodha's trading database processes 100,000 transactions per second at market open. Choosing gp2 instead of io2 for the database volume caused IOPS throttling at peak — transaction latencies spiked and the platform degraded under load. The fix was a volume type change and decoupling IOPS from storage size using gp3. At Swiggy, a developer accidentally deleted a critical snapshot used for production database restores. Without Recycle Bin it would have been gone permanently. Storage decisions feel invisible until they cause an incident — at which point the impact is immediate and production-wide.
What is Amazon EBS
EBS (Elastic Block Store) is network-attached block storage for EC2 instances. Think of it as a USB hard drive that attaches over a network — you can detach it from one instance and attach it to another.
Characteristics: Persists data independently of the EC2 instance lifecycle Can be detached and reattached to another EC2 instance Locked to one Availability Zone — cannot cross AZs without snapshot Provisioned capacity — you pay for what you allocate, not what you use One EBS volume → one EC2 instance at a time (except io1/io2 Multi-Attach) What it is NOT: Not a file system (you must format it — ext4, xfs, ntfs) Not shared across instances by default (use EFS for shared access) Not the fastest possible storage (Instance Store is faster)Delete on Termination:
When you launch an EC2 instance, the root EBS volume has Delete on Termination enabled by default. Additional volumes have it disabled by default.
Root volume (Delete on Termination = true by default): Instance terminated → root volume deleted automatically Additional data volumes (Delete on Termination = false by default): Instance terminated → data volume persists until you delete it manuallyRememberDisable Delete on Termination on root volumes for production databases and any instance where you need to recover the disk after an accidental termination.
EBS Volume Types — The Full Breakdown
Six volume types organised into three categories:
SSD-backed (IOPS-intensive): gp3 — General Purpose SSD (new, always prefer over gp2) gp2 — General Purpose SSD (old, being phased out) io1 — Provisioned IOPS SSD (legacy high performance) io2 Block Express — Provisioned IOPS SSD (current high performance) HDD-backed (Throughput-intensive): st1 — Throughput Optimized HDD sc1 — Cold HDDRememberOnly SSD volumes (gp2, gp3, io1, io2) can be used as boot volumes. HDD volumes (st1, sc1) cannot be boot volumes.
gp3 vs gp2 — The Most Important Distinction
This is the most tested EBS concept. gp3 replaced gp2 as the default general-purpose volume.
gp2 — old, IOPS tied to size:
IOPS = 3 × storage GB (with a baseline of 100 IOPS minimum)1000 GB gp2 → 3000 IOPS (fixed, cannot change independently)Want 6000 IOPS? You must provision 2000 GB even if you need only 100 GB of spaceYou pay for unnecessary storage to get the IOPS you need Burst: below 1000 GB instances burst to 3000 IOPS using I/O creditsMax: 16,000 IOPSgp3 — new, IOPS and throughput independent of size:
Baseline: 3000 IOPS and 125 MB/s included with every volume regardless of sizeYou can add up to 16,000 IOPS and 1000 MB/s independently of storage GB20% cheaper than gp2 for the same storage size gp3 with 100 GB → still gets 3000 IOPS (no need to over-provision storage)gp3 with 100 GB → can scale to 16,000 IOPS by paying for IOPS separatelyThe critical difference:
gp2: MORE STORAGE = MORE IOPS (coupled, cannot separate)gp3: STORAGE AND IOPS ARE INDEPENDENT (decoupled, scale each separately) Example — you need 6000 IOPS and 100 GB storage: gp2: must provision 2000 GB to get 6000 IOPS → 1900 GB wasted gp3: provision 100 GB, add 6000 IOPS separately → no wasteRemembergp3 decouples IOPS from storage. gp2 ties them together. For any new volume, always use gp3 — it is cheaper and more flexible. Migrate existing gp2 volumes to gp3.
io1 and io2 — Provisioned IOPS for Critical Databases
When gp3's maximum 16,000 IOPS is not enough, use Provisioned IOPS SSD.
io1 (legacy):
Max IOPS: 64,000 (on Nitro instances)Max throughput: 1,000 MB/sIOPS to storage ratio: max 50:1 (1000 GB → max 50,000 IOPS)Multi-Attach: yesio2 Block Express (current):
Max IOPS: 256,000Max throughput: 4,000 MB/sIOPS to storage ratio: max 1000:1 (much more efficient)Sub-millisecond latencyMulti-Attach: yesUse for: SAP HANA, Oracle, Microsoft SQL Server, large PostgreSQLEBS Multi-Attach:
io1 and io2 support attaching one volume to multiple EC2 instances in the same AZ simultaneously (up to 16 instances).
Use case: Cluster-aware applications that manage concurrent writes themselves Must use a filesystem that supports concurrent access (ex4 does NOT) Example: Lustre on HPC workloads, Oracle RACCommon MistakeUsing Multi-Attach with standard filesystems like ext4. Standard filesystems are not designed for multi-host concurrent access. Data corruption occurs. Only use Multi-Attach with cluster-aware applications that handle concurrency at the application level.
st1 and sc1 — HDD Volumes for Throughput Workloads
HDD volumes are for sequential read/write workloads where throughput matters more than IOPS. Much cheaper than SSD per GB.
st1 — Throughput Optimized HDD:
Max throughput: 500 MB/sMax IOPS: 500Cannot be boot volumeUse for: big data (Hadoop, Kafka), data warehouses, log processing, ETL Pattern: reading massive files sequentially end to endWrong for: databases with random access patternssc1 — Cold HDD:
Max throughput: 250 MB/sMax IOPS: 250Cheapest EBS optionCannot be boot volumeUse for: infrequently accessed data, archives, compliance storageEBS Type Summary:
| Type | Max IOPS | Max Throughput | Boot? | Best For |
|---|---|---|---|---|
| gp3 | 16,000 | 1,000 MB/s | Yes | Default for everything |
| gp2 | 16,000 | 250 MB/s | Yes | Legacy only |
| io2 Block Express | 256,000 | 4,000 MB/s | Yes | Mission-critical databases |
| io1 | 64,000 | 1,000 MB/s | Yes | Legacy high-performance |
| st1 | 500 | 500 MB/s | No | Big data, streaming logs |
| sc1 | 250 | 250 MB/s | No | Cold archival storage |
EBS Encryption
EBS encryption uses AWS KMS. When you create an encrypted volume:
Data at rest encrypted on the volumeData in-flight between EC2 and volume encryptedAll snapshots encryptedAll volumes created from encrypted snapshots are encryptedFour steps to encrypt an existing unencrypted volume:
This is the most commonly tested EBS workflow.
Step 1: Create an EBS Snapshot of the unencrypted volumeStep 2: Copy the snapshot — enable encryption during the copy (choose KMS key)Step 3: Create a new EBS volume from the encrypted snapshot copyStep 4: Attach the new encrypted volume to the instance, detach the old unencrypted one## Step 1 — Snapshot the unencrypted volumeaws ec2 create-snapshot \ --volume-id vol-0abc123def456789 \ --description "Pre-encryption snapshot" \ --region ap-south-1 ## Step 2 — Copy snapshot with encryption enabledaws ec2 copy-snapshot \ --source-region ap-south-1 \ --source-snapshot-id snap-0abc123456789def \ --description "Encrypted copy" \ --encrypted \ --kms-key-id alias/devops-prod-key \ --destination-region ap-south-1 ## Step 3 — Create encrypted volume from the encrypted snapshotaws ec2 create-volume \ --snapshot-id snap-0encrypted123456 \ --volume-type gp3 \ --availability-zone ap-south-1a \ --region ap-south-1 ## Step 4 — Detach old, attach newaws ec2 detach-volume \ --volume-id vol-0abc123def456789 \ --region ap-south-1 aws ec2 attach-volume \ --volume-id vol-0newencryptedvol \ --instance-id i-0abc123def456789 \ --device /dev/sdf \ --region ap-south-1TipEnable account-level EBS encryption by default in every region. New volumes are automatically encrypted with the default KMS key. Zero effort, zero risk. Setting → EC2 → EBS Encryption → Enable.
EBS Snapshots
Snapshots are point-in-time backups of EBS volumes stored in S3. They are incremental — only changed blocks since the last snapshot are saved.
First snapshot: copies all 50 GB of data → takes timeSecond snapshot (after 2 GB changed): copies only 2 GB → fast, cheap You can restore any snapshot regardless of orderYou can delete intermediate snapshots — the restore process reconstructs themCross-region copy and cross-account sharing:
## Copy snapshot to another region (for DR)aws ec2 copy-snapshot \ --source-region ap-south-1 \ --source-snapshot-id snap-0abc123456789def \ --destination-region ap-southeast-1 \ --description "DR copy to Singapore" \ --region ap-southeast-1 ## Share snapshot with another AWS accountaws ec2 modify-snapshot-attribute \ --snapshot-id snap-0abc123456789def \ --attribute createVolumePermission \ --operation-type add \ --user-ids 999888777666 \ --region ap-south-1EBS Snapshot Archive:
Move a snapshot to Archive tier for 75% lower storage cost. Restoring from Archive takes 24-72 hours — use only for snapshots you rarely need to restore.
## Move snapshot to Archive tieraws ec2 modify-snapshot-tier \ --snapshot-id snap-0abc123456789def \ --storage-tier archive \ --region ap-south-1Recycle Bin — Snapshot Recovery
Without Recycle Bin, a deleted snapshot is permanently gone immediately. No recovery.
Recycle Bin holds deleted snapshots for a retention period you define before they are permanently destroyed.
## Create a Recycle Bin retention rule — keep deleted snapshots for 14 daysaws rbin create-rule \ --retention-period RetentionPeriodValue=14,RetentionPeriodUnit=DAYS \ --resource-type EBS_SNAPSHOT \ --description "14 day snapshot recovery window" \ --region ap-south-1 ## List snapshots currently in Recycle Binaws rbin list-resources \ --resource-type EBS_SNAPSHOT \ --region ap-south-1 ## Recover a specific snapshot from Recycle Binaws rbin recover-resource \ --resource-id snap-0abc123456789def \ --resource-type EBS_SNAPSHOT \ --region ap-south-1TipEnable Recycle Bin for EBS snapshots in every production account from day one. A 7 to 30 day retention rule costs almost nothing and has saved teams from catastrophic data loss after accidental deletes.
Fast Snapshot Restore (FSR)
When you create an EBS volume from a snapshot, the volume starts empty and data is loaded lazily from S3 as you access each block for the first time. This causes high initial latency — the first access to any block is slow.
This is called initialisation latency and it matters for databases being restored from snapshots.
Without FSR: Restore DB from snapshot → first query hits cold blocks → slow response Initialisation takes hours of I/O before full performance With FSR enabled on a snapshot: AWS pre-warms the snapshot in specific AZs Volume created from FSR snapshot → immediately full performance No initialisation period## Enable FSR on a specific snapshot in ap-south-1aaws ec2 enable-fast-snapshot-restores \ --availability-zones ap-south-1a ap-south-1b \ --source-snapshot-ids snap-0abc123456789def \ --region ap-south-1FSR is billed per minute per snapshot per AZ. Enable it only for snapshots used in time-sensitive restores.
EC2 Instance Store — Ephemeral Local Storage
While EBS is a network drive, Instance Store is physical storage directly attached to the host server where your EC2 instance runs.
EBS: Network-attached → some latency Persists after instance stop/terminate Can be detached and reattached Limited by network bandwidth Instance Store: Physically attached → zero network latency EPHEMERAL — data lost when instance stops or terminates Cannot be detached — tied to the physical host Fastest possible I/O on EC2Instance Store performance vs EBS:
| Volume | Max IOPS | Max Throughput | Persists? |
|---|---|---|---|
| io2 Block Express | 256,000 | 4,000 MB/s | Yes |
| Instance Store (i4i.32xlarge) | 4,000,000+ | Very high | NO |
Instance Store IOPS can be 10-20x higher than even io2 Block Express.
When Instance Store is the correct choice:
Temporary data that can be lost and recreated: Buffer and cache data (warm the cache from the database on restart) Scratch space for ML training or data processing Replicated data (each node has a copy, one node failing is acceptable) Applications designed for ephemeral storage: Kafka brokers (data replicated across the cluster) Elasticsearch nodes (data replicated across the cluster) HPC scratch storageCommon MistakeUsing Instance Store for data that must persist. An instance stop, hardware failure, or termination destroys all Instance Store data permanently — no snapshot, no recovery. Never use Instance Store for a production database unless data is replicated across multiple nodes and the application can rebuild from replicas.
Hands-on Lab — Create EBS Volume, Snapshot, and Encrypt
Step 1 — Create and attach a new EBS volume
EC2 → Volumes → Create volumeVolume type: gp3Size: 10 GiBAvailability Zone: same AZ as your EC2 instance (must match)IOPS: 3000 Throughput: 125 MB/s (both independent in gp3 — this is the gp3 advantage)Create volume Select the new volume → Actions → Attach volumeInstance: select your running EC2 → Device: /dev/xvdf → AttachStep 2 — Format and mount the volume
SSH into your EC2 instance:## See the new disklsblk## xvdf appears — unformatted ## Format with ext4sudo mkfs -t ext4 /dev/xvdf ## Create mount point and mountsudo mkdir /datasudo mount /dev/xvdf /data ## Verifydf -h## /dev/xvdf shows 10G availableStep 3 — Write data and take a snapshot
## Write test data to the volumeecho "Important production data" | sudo tee /data/critical.txtcat /data/critical.txtEC2 → Volumes → select your gp3 volumeActions → Create snapshotDescription: test-snapshot-before-changeCreate snapshot EC2 → Snapshots → watch status change from pending → completedStep 4 — Restore from snapshot
EC2 → Snapshots → select your snapshotActions → Create volume from snapshotVolume type: gp3 Size: 10 GiBAZ: any AZ you want (snapshots can restore to any AZ)Create volume The new volume contains exactly the data that existed at snapshot time.Step 5 — Encrypt an unencrypted volume
EC2 → Snapshots → select your snapshotActions → Copy snapshotEncryption: EnableKMS key: AWS managed key (aws/ebs)Copy snapshot EC2 → Snapshots → new encrypted snapshot createdActions → Create volume from snapshot → new encrypted volume Attach encrypted volume to EC2 → mount it → data is the sameVolume is now encrypted at rest.Step 6 — Test Recycle Bin for snapshot recovery
EC2 → Recycle Bin → Retention rules → Create retention ruleResource type: EBS snapshotsRetention period: 7 days → Create Now delete one of your test snapshots.EC2 → Recycle Bin → Resources → find your deleted snapshotActions → Recover → snapshot is restored and usable again.Step 7 — Cleanup
Unmount volume: sudo umount /dataEC2 → Volumes → detach and delete test volumesEC2 → Snapshots → delete test snapshotsRecycle Bin → delete retention ruleProduction Best Practices and Common Pitfalls
- Use gp3 for all new volumes — it is cheaper than gp2 and decouples IOPS from storage size
- Enable account-level EBS encryption by default — zero effort, all new volumes automatically encrypted
- Enable Recycle Bin in every production account with at least a 7-day retention — costs almost nothing and protects against accidental deletes
- Enable Delete on Termination = false for data volumes — never allow auto-deletion of volumes holding data on instance termination
- Take snapshots before any risky operation — resize, type change, AMI migration
- Use st1 for sequential big-data workloads instead of gp3 — st1 is significantly cheaper for high-throughput sequential access
- Never use Instance Store for data that must survive an instance stop — only for ephemeral buffers and replicated data
Quick Reference and Troubleshooting Commands
| Task | Command |
|---|---|
| List volumes | aws ec2 describe-volumes --region ap-south-1 |
| Create volume | aws ec2 create-volume --size 20 --volume-type gp3 --availability-zone ap-south-1a --encrypted |
| Attach volume | aws ec2 attach-volume --volume-id <vol-id> --instance-id <i-id> --device /dev/sdf |
| Detach volume | aws ec2 detach-volume --volume-id <vol-id> --region ap-south-1 |
| Create snapshot | aws ec2 create-snapshot --volume-id <vol-id> --description "backup" |
| List snapshots | aws ec2 describe-snapshots --owner-ids self --region ap-south-1 |
| Copy snapshot to another region | aws ec2 copy-snapshot --source-region ap-south-1 --source-snapshot-id <snap-id> --destination-region ap-southeast-1 |
| Modify volume type | aws ec2 modify-volume --volume-id <vol-id> --volume-type gp3 --iops 6000 |
| Check modification status | aws ec2 describe-volumes-modifications --volume-ids <vol-id> |
Common problems and fixes:
| Problem | Likely cause | Fix |
|---|---|---|
| Volume in wrong AZ | EBS is AZ-locked | Snapshot the volume, copy to correct AZ, create new volume |
| First reads slow after snapshot restore | Initialisation latency — blocks loaded lazily | Enable Fast Snapshot Restore or pre-warm by reading all blocks |
| IOPS throttling on gp2 at peak | IOPS tied to size, not enough storage to get needed IOPS | Migrate to gp3, set IOPS independently |
| Volume not visible after attach | Device path mismatch or kernel naming | Check lsblk inside EC2, use NVMe naming for Nitro instances |
| Cannot delete snapshot | Snapshot referenced by an AMI | Deregister the AMI first, then delete the snapshot |
Common MistakeMigrating gp2 to gp3 without checking current IOPS usage. gp2 provides burst IOPS that gp3 does not. A gp2 volume with burst credit balance may be running above its baseline. After migrating to gp3, set explicit IOPS at least equal to what the workload was actually using — not just the gp2 baseline.
TipYou can change volume type, size, and IOPS without any downtime on a running instance using Modify Volume. The modification completes in the background while your instance keeps running. Monitor the modification progress with describe-volumes-modifications until state = completed before making another change to the same volume.
Common Mistakes to Avoid
Common MistakeUsing gp2 when gp3 is available. gp3 is cheaper, faster, and decouples IOPS from storage size. There is no reason to use gp2 for any new volume.
Common MistakeUsing Instance Store for data that must survive. Instance Store is ephemeral — stop or terminate the instance and all data is permanently gone. Never store anything on Instance Store that you cannot recreate instantly.
RememberEBS snapshots are incremental. Only changed blocks since the last snapshot are stored. Delete intermediate snapshots freely — each remaining snapshot is still self-contained and fully restorable.