Skip to main content

Amazon S3

Simple Storage Service. Object storage for files of any type up to 5 TB each. Globally unique bucket names, virtually unlimited capacity, and 11 nines of durability across multiple Availability Zones.

S3 (Simple Storage Service) is AWS object storage. You store files — called objects — inside containers called buckets. No capacity to provision, no filesystem to manage. Virtually unlimited scale with 11 nines of durability.

Objects and Keys

An object is any file. The key is its full path inside the bucket — including what looks like folder names (they are just slashes in the key, not real directories).

TEXT
s3://hotstar-assets-prod/thumbnails/shows/mirzapur/ep1.jpg
Bucket: hotstar-assets-prod
Key: thumbnails/shows/mirzapur/ep1.jpg
Prefix: thumbnails/shows/mirzapur/

Storage Classes — Automatic Cost Optimisation

Class Access pattern Retrieval fee Min duration Use for
Standard Frequently accessed None None Active data
Standard-IA Infrequent, needs fast retrieval Per GB 30 days Backups, DR copies
One Zone-IA Infrequent, can be recreated Per GB 30 days Derived data like thumbnails
Glacier Instant Quarterly access Per GB 90 days Archives needing ms retrieval
Glacier Flexible Annual access Per GB 90 days Long-term archives
Deep Archive 7-10 year retention Per GB 180 days Regulatory archives
Intelligent-Tiering Unknown pattern None None When access pattern is unpredictable

Lifecycle rules automate transitions: Standard → Standard-IA after 30 days → Glacier after 90 days → expire after 365 days.

Access Control — The Decision Rule

A request to S3 is allowed if: (IAM policy allows it) OR (Bucket policy allows it) AND (no explicit Deny anywhere).

An explicit Deny always wins — over any Allow, in any policy, anywhere.

Block Public Access

Four settings that override bucket policies and ACLs trying to make buckets public. All four are ON by default for every new bucket. This is the safety net that prevents accidental data exposure.

Only disable Block Public Access when you specifically need a public bucket (static website, public downloads). Every other bucket should have all four settings on.

Versioning

◈ DIAGRAM
Without versioning: overwrite a file → old version permanently gone
With versioning: overwrite a file → new version ID assigned, old version preserved
Delete with versioning: creates a delete marker — file appears gone, versions still exist
Restore: delete the delete marker → original file reappears

Performance

◈ DIAGRAM
3,500 PUT/COPY/POST/DELETE requests per second per prefix
5,500 GET/HEAD requests per second per prefix
Spread objects across many prefixes to multiply throughput
Files above 5 GB: must use multipart upload
Files above 100 MB: recommended to use multipart upload
Parallel part uploads → faster, failed part = retry only that part

Encryption

TEXT
SSE-S3: default, AWS manages keys, zero config
SSE-KMS: KMS manages keys, full CloudTrail audit trail per access
SSE-C: you provide keys in request headers
Client-side: you encrypt before upload, AWS never sees plaintext
Remember

Bucket names are globally unique across all AWS accounts worldwide. A name used by anyone is unavailable to everyone else. Use a pattern like company-project-environment-accountid to avoid conflicts.

Common Mistake

Setting the S3 server access log destination to the same bucket being logged. Every access creates a log → which is an access → which creates another log → bucket grows uncontrollably. Always use a separate dedicated logging bucket.

Frequently Asked Questions

What does 'eleven nines of durability' actually mean for S3, and how is that different from availability?

Durability (99.999999999%) is about not losing your data — S3 replicates every object across multiple devices in multiple Availability Zones, so the odds of permanent data loss are vanishingly small. Availability (99.9% typical SLA) is about whether you can access it right now — a network blip or regional issue could make S3 briefly unreachable without any data being lost. They're separate guarantees people often conflate.

What's a common mistake engineers make with S3 in production?

Leaving buckets with default public-access settings misconfigured, which has caused numerous real-world data leaks. Also common: not setting a lifecycle policy, so objects sit in Standard storage indefinitely instead of transitioning to Infrequent Access or Glacier as they age, quietly inflating storage bills. Enable Block Public Access by default and pair it with lifecycle rules tuned to actual access patterns.