What you will learn
- What S3 is and how it is different from a file system or a database
- Bucket naming rules and how objects are actually stored
- Bucket policies vs IAM policies — how S3 decides who gets access
- Block Public Access — the safety switch that overrides everything
- Versioning — how delete markers work and how to recover deleted files
- Replication — CRR and SRR, what gets replicated and what does not
- All storage classes and how lifecycle rules move objects between them automatically
- S3 performance — multipart upload, Transfer Acceleration, and byte-range fetches
- Encryption — four methods and when to use each
- Pre-signed URLs for time-limited access to private files
Why this matters
Hotstar serves video thumbnails and metadata to millions of users through S3 and CloudFront. Razorpay stores payment compliance documents with Object Lock so they cannot be modified or deleted. Swiggy aggregates application logs from hundreds of services into a single S3 bucket and queries them with Athena. Zerodha stores trading records for regulatory retention requirements using Glacier Deep Archive at a fraction of the cost of a database.
S3 is not just file storage. It is the integration point for almost every AWS service. And it is involved in one of the most common AWS security incidents — accidentally public buckets containing sensitive data. Understanding S3 properly means understanding both how powerful it is and where the traps are.
What is Amazon S3
S3 (Simple Storage Service) is object storage. You store files — called objects — inside containers called buckets. Unlike a file system, there are no real folders. Unlike a database, there is no schema or query language.
Object storage vs file system: File system: /var/app/uploads/user123/photo.jpg (real directory tree) S3: s3://my-bucket/uploads/user123/photo.jpg (flat, key is the path) In S3 what looks like a folder is just a prefix in the object key name.There are no actual directories. The slash is part of the filename.S3 is accessed via HTTP/HTTPS using an API — not mounted like a disk. Your application calls the S3 API to put, get, or delete objects.
What makes S3 special:
Virtually unlimited storage — no capacity to provision11 nines of durability — 99.999999999% — data stored across multiple AZsGlobally unique bucket names — no two buckets in the world share a nameObjects up to 5 TB eachPay only for what you store — no minimumBuckets and Bucket Naming Rules
A bucket is the top-level container. Every object lives inside a bucket.
Naming rules — every one matters:
Must be globally unique across all AWS accounts worldwide3 to 63 characters longLowercase letters, numbers, and hyphens only — no uppercase, no underscoresMust start with a letter or numberCannot look like an IP address (not 192.168.1.1)Cannot start with xn-- Valid: razorpay-payment-logs-2024Valid: hotstar-thumbnails-prodInvalid: Hotstar-Thumbnails (uppercase H)Invalid: my_bucket (underscore not allowed)Invalid: ab (too short)Buckets are created in a specific region. Your data stays in that region unless you explicitly replicate it. Despite looking global in the console, every bucket has a home region.
Objects and Keys
An object is any file stored in S3. The key is the full path to that object inside the bucket.
Bucket: razorpay-payment-logsKey: invoices/2024/january/INV-001.pdf Full address: s3://razorpay-payment-logs/invoices/2024/january/INV-001.pdfThe key has two parts:
Prefix: invoices/2024/january/Object name: INV-001.pdfRememberThere are no real folders in S3. What looks like a folder hierarchy is just slashes in the key name. The entire path is the filename. This matters for performance — spreading objects across many different prefixes dramatically increases read and write throughput.
Object size limits:
| Situation | Limit |
|---|---|
| Maximum object size | 5 TB |
| Maximum single upload (PUT) | 5 GB |
| Files above 5 GB | Must use multipart upload |
| Recommended multipart threshold | 100 MB |
S3 Security — Who Gets Access
S3 access is controlled at two levels and can come from two directions.
The access decision rule:
A request to S3 is ALLOWED if: The IAM policy of the caller ALLOWS it OR the Bucket Policy ALLOWS it AND there is no explicit DENY anywhere An explicit DENY always wins — no exceptionsBucket Policy — the resource-side rule:
A bucket policy is a JSON document attached directly to the bucket. It controls access to that bucket for anyone — IAM users, other accounts, anonymous public users, AWS services.
Scenario 1 — Make a bucket public (static website): No IAM user — anonymous public access needed Solution: Bucket policy with Principal: "*" allows s3:GetObject Scenario 2 — EC2 accessing S3: IAM Role attached to EC2 with S3 read permission No bucket policy needed — IAM role handles it Scenario 3 — Cross-account access: Bucket policy allows the other account's ARN as Principal{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": "*", "Action": "s3:GetObject", "Resource": "arn:aws:s3:::my-static-website/*" } ]}Remember
arn:aws:s3:::my-buckettargets the bucket itself (for listing).arn:aws:s3:::my-bucket/*targets objects inside the bucket (for reading files). Many access problems come from using the wrong ARN. Listing needs the bucket ARN. Getting objects needs the/*ARN. Often you need both statements.
Block Public Access
Block Public Access is a safety override that ignores bucket policies and ACLs that try to make your bucket public. It was created after too many data breaches from accidentally public S3 buckets.
Four settings — each can be on or off individually:
Block public access granted through new ACLsBlock public access granted through any ACLsBlock public access granted through new bucket policiesBlock public and cross-account access through any policiesDefault: All four are ON for every new bucket. AWS also applies them at the account level.
Even if a bucket policy explicitly says "allow public access", Block Public Access overrides it and blocks anyway. This is the safety net.
Only turn off Block Public Access when you specifically need a public bucket — static website hosting, public file downloads. For every other bucket keep all four settings on.
Versioning
Versioning keeps every version of every object. Once enabled, uploading the same key creates a new version — it does not overwrite the old one.
Without versioning: Upload report.pdf → stored Upload report.pdf again → old version permanently gone With versioning: Upload report.pdf → version ID: aaa111 Upload report.pdf again → version ID: bbb222 (both exist) Upload report.pdf again → version ID: ccc333 (all three exist)Key facts:
Files uploaded before versioning was enabled → version ID is nullSuspending versioning → does not delete existing versionsSuspending versioning → stops creating new versions going forwardHow deletes work with versioning:
Without versioning: Delete report.pdf → permanently gone With versioning: Delete report.pdf → S3 creates a delete marker The file looks gone in the console But all previous versions still exist To restore: delete the delete marker → file reappears Permanent delete: Delete a specific version by its version ID → that version is gone foreverTipEnable versioning on any bucket that holds data you cannot afford to lose. It protects against two things: accidental overwrites (restore a previous version) and accidental deletes (remove the delete marker). The extra storage cost is usually worth it.
Replication — CRR and SRR
Replication automatically copies new objects from one bucket to another in the background.
Cross-Region Replication (CRR):
Source bucket in ap-south-1 → replicated to destination in us-east-1Use for: disaster recovery, lower latency for global users, cross-account backupSame-Region Replication (SRR):
Source bucket in ap-south-1 → replicated to another bucket in ap-south-1Use for: combining logs from many buckets into one, live copy between prod and testRequirements for either type:
Versioning must be enabled on BOTH source and destination bucketProper IAM permissions for S3 to perform the replicationWhat replication does and does not do:
New objects after enabling → replicated automaticallyExisting objects before enabling → NOT replicated (use S3 Batch Replication) Delete markers → can be replicated (optional, you choose)Deleting a specific version ID → NEVER replicatedThis protects you: a malicious delete cannot propagate to your backup bucket No chaining:Bucket A replicates to Bucket BBucket B does NOT automatically replicate to Bucket CIf you need A → C, create a direct replication rule from A to CStorage Classes — The Cost vs Speed Tradeoff
All S3 storage classes have the same 11 nines of durability. What changes is availability, retrieval speed, and cost.
| Storage Class | Access Pattern | Min Duration | Retrieval | Best For |
|---|---|---|---|---|
| S3 Standard | Frequently accessed | None | Free | Active data, production files |
| S3 Standard-IA | Infrequent but fast needed | 30 days | Per GB | Backups, DR copies |
| S3 One Zone-IA | Infrequent, can recreate | 30 days | Per GB | Thumbnails, derived files |
| S3 Glacier Instant | Quarterly access | 90 days | Per GB | Archives needing ms retrieval |
| S3 Glacier Flexible | Annual access | 90 days | Per GB | Long-term archives (hours OK) |
| S3 Glacier Deep Archive | 7-10 year retention | 180 days | Per GB | Regulatory archives |
| S3 Intelligent-Tiering | Unknown pattern | None | None | When access pattern is unpredictable |
The minimum billing duration trap:
Glacier Flexible: minimum 90 days billingStore a file for 1 day then delete → still billed for 90 daysDeep Archive: minimum 180 daysThese are for data you are keeping long-term, not short-term filesIntelligent-Tiering — fully automated:
Monitors access for each object. Moves it between tiers automatically based on actual usage patterns. No retrieval fees. Small monitoring fee per object. Use when you genuinely do not know how often data will be accessed.
Lifecycle Rules — Automatic Cost Reduction
A lifecycle rule automatically transitions objects between storage classes after a set number of days, and optionally deletes them after another set of days.
Example — application logs: Day 0: Logs uploaded to Standard (active investigation) Day 30: Move to Standard-IA (still occasionally referenced) Day 90: Move to Glacier Flexible (archived, rarely touched) Day 365: Delete (retention period met) One lifecycle rule. Zero human action after setup.Rules can target everything in the bucket or just specific prefixes:
Prefix: logs/ → only log files get this rulePrefix: uploads/ → only user uploads get this ruleRememberAlways add a lifecycle rule to delete incomplete multipart uploads after 7 days. A failed large upload leaves behind parts that bill you but never become a complete object. Without this rule those orphaned parts accumulate silently.
S3 Performance
S3 scales automatically but some patterns can squeeze much more throughput out of it.
Baseline performance:
3,500 PUT/COPY/POST/DELETE requests per second per prefix5,500 GET/HEAD requests per second per prefixSpread objects across many prefixes and the numbers multiply:
1 prefix → 5,500 GET/sec4 prefixes → 22,000 GET/sec10 prefixes → 55,000 GET/secMultipart Upload — for large files:
Breaks a large file into parts and uploads all parts in parallel. Faster. If one part fails, retry only that part, not the entire file.
Required above 5 GBRecommended above 100 MBAll parts upload simultaneously → much faster than sequentialS3 assembles the parts into the complete object automaticallyS3 Transfer Acceleration:
Routes uploads through the nearest AWS Edge Location. Short fast hop to the edge, then AWS's fast private network the rest of the way to your bucket.
Without: Your office in Chennai → slow public internet → S3 in ap-south-1With: Your office → Mumbai Edge → AWS private network → S3 bucketUseful when uploading from distant locations or across continents. Small per-GB cost.
Byte-Range Fetches — download only what you need:
Request a specific byte range of an object instead of downloading the entire file.
Read only the first 1 KB of a 500 MB file: Without: download 500 MB → read 1 KB → discard 499 MB With: request bytes 0-1023 → receive 1 KB → done Download a large file in parallel: Request bytes 0-999 simultaneously with bytes 1000-1999 with bytes 2000-2999 All three download at the same time → 3x fasterEncryption
Four methods. The right one depends on who should control the keys and your compliance requirements.
| Method | Who manages keys | HTTPS required | Default now |
|---|---|---|---|
| SSE-S3 | AWS fully | No | Yes — all new buckets |
| SSE-KMS | AWS KMS (you control) | No | No |
| SSE-C | You (outside AWS) | Yes — mandatory | No |
| Client-Side | You fully | No | No |
SSE-S3 — default, zero effort:
AWS encrypts every object with AES-256. You do nothing. Best for most workloads that just need encryption without special compliance requirements.
SSE-KMS — full audit trail:
Keys managed in AWS KMS. Every encryption and decryption operation appears in CloudTrail. Use for compliance requirements where you need proof of who accessed encrypted data and when.
One trap: every S3 upload with SSE-KMS calls KMS GenerateDataKey. Every download calls KMS Decrypt. KMS has API quotas. At high request rates (millions of S3 operations per hour) you can hit the KMS quota and get throttled. If this happens, request a quota increase or switch to SSE-S3.
SSE-C — keys you manage, AWS never stores:
You send your encryption key in the request header. AWS encrypts, then immediately forgets your key. HTTPS is mandatory because your key travels in the header. AWS stores the encrypted data, not your key.
Client-Side — AWS never sees plaintext:
You encrypt on your machine before uploading. S3 stores an already-encrypted blob. S3 has no idea what is inside. Only clients with the correct key can decrypt. Maximum control.
Pre-Signed URLs
A pre-signed URL gives temporary access to a private S3 object without making the bucket public.
You generate a URL using your credentialsThe URL contains: the object path + operation + expiry time + your signatureYou share the URL with anyoneThey use it to download (or upload) the file directly from S3The URL expires automatically after the set time## Generate a URL valid for 1 hour (3600 seconds)aws s3 presign s3://razorpay-invoices/INV-2024-001.pdf \ --expires-in 3600 \ --region ap-south-1## Output is a long URL anyone can use to download the file until it expiresUse cases:
Let a logged-in user download their invoice — generate a URL per sessionLet a specific user upload a file without giving them permanent S3 accessShare a private document with a client for 24 hours then access automatically revokesS3 Event Notifications — React to What Happens in Your Bucket
Every time an object is uploaded, deleted, or restored, S3 can automatically notify other services. No polling needed. No cron jobs. The bucket tells the world something happened.
User uploads a photo to S3 ↓S3 fires an event notification immediately ↓Lambda triggered → resizes photo, creates thumbnailSQS queue receives message → worker picks it up for virus scanningSNS topic notified → email alert sent to ops teamFour destinations for event notifications:
| Destination | Best for |
|---|---|
| Lambda | Immediate serverless processing |
| SQS | Queued processing by workers |
| SNS | Fan-out to multiple subscribers |
| EventBridge | Complex routing, filtering, and cross-account events |
EventBridge is the most powerful option — it lets you filter events by prefix, suffix, or object metadata and route them to dozens of different targets including Step Functions, Kinesis, and other AWS accounts.
Example — Hotstar video upload pipeline:Video uploaded to s3://hotstar-uploads/raw/ ↓S3 Event → SQS queue ↓EC2 worker picks up message → transcodes video → saves to s3://hotstar-cdn/One important rule: S3 sends each event notification at least once. Your consumer must handle duplicates.
S3 Object Lock and Glacier Vault Lock — WORM Storage
WORM stands for Write Once Read Many. Once written, the object cannot be modified or deleted — by anyone, including the root user — until the retention period expires.
Razorpay uses Object Lock on compliance documents that regulators require to be kept unmodified for 7 years. Nobody can delete or overwrite them accidentally or maliciously.
Two Object Lock modes:
Governance Mode: Most users cannot delete or modify Users with special IAM permissions CAN override the lock Use for: internal compliance where occasional exceptions are needed Compliance Mode: Nobody can delete or modify — including root user Retention period cannot be shortened once set Use for: regulatory requirements where no exceptions are allowedRetention Period vs Legal Hold:
Retention Period: lock expires after N days automaticallyLegal Hold: lock stays until you explicitly remove it — no expiry dateBoth can be on the same object simultaneouslyGlacier Vault Lock:
Same concept but for Glacier Vaults. Once a Vault Lock policy is locked (confirmed), it cannot be changed or deleted. Complete regulatory compliance for archived data.
RememberObject Lock must be enabled when the bucket is created. You cannot enable it on an existing bucket. Versioning is automatically enabled when Object Lock is enabled. Plan for this before creating compliance buckets.
S3 Access Points — Simplify Multi-Team Access
A large company has one S3 bucket with data for many teams — Finance, Engineering, Analytics, Compliance. Managing bucket policies that correctly grant each team access to only their data gets complicated fast.
S3 Access Points solve this. Each team gets their own Access Point — their own hostname and their own policy — pointing to the same underlying bucket.
s3://company-data-lake/ (one bucket, all data)├── finance-ap.s3-accesspoint... → Finance team, /finance/* prefix only├── engineering-ap.s3-accesspoint... → Engineering, /engineering/* only└── analytics-ap.s3-accesspoint... → Analytics, read-only on everythingEach Access Point has its own policy. Adding a new team means creating a new Access Point — the bucket policy stays simple and unchanged.
Object Lambda — Transform Data on the Way Out:
Object Lambda sits between S3 and the application. When an object is retrieved through an Object Lambda Access Point, a Lambda function runs and can modify the response before it reaches the caller.
Application requests customer data from S3 ↓Object Lambda intercepts ↓Lambda function redacts PII (phone numbers, email addresses) ↓Application receives sanitised dataSame underlying data. Different views for different teams. No duplicate storage needed.
Static Website Hosting and CORS
S3 can serve a static website — HTML, CSS, JavaScript files — directly. No web server. No EC2.
S3 → Bucket → Properties → Static website hosting → EnableSet index document: index.htmlSet error document: error.html (optional)S3 gives you a website endpoint URLFor a public website the bucket must allow public read access. Combine with CloudFront to add HTTPS, caching, and a custom domain.
CORS — Cross-Origin Resource Sharing:
If your website at https://myapp.in makes API calls to https://api.myapp.in, the browser blocks the request unless the API explicitly says the other origin is allowed. This is CORS.
When your web app hosted on one domain needs to fetch assets from an S3 bucket on another domain, S3 needs a CORS configuration.
S3 → Bucket → Permissions → Cross-origin resource sharing (CORS)[ { "AllowedOrigins": ["https://myapp.in"], "AllowedMethods": ["GET"], "AllowedHeaders": ["*"], "MaxAgeSeconds": 3600 }]This tells browsers: requests from https://myapp.in are allowed to fetch objects from this bucket.
Hands-on Lab — Versioning, Lifecycle, and Event Notifications
Step 1 — Create a versioned bucket
S3 → Create bucketName: devops-lab-2024-yourname (must be globally unique)Region: ap-south-1Block Public Access: keep all ONVersioning: Enable → Create bucketStep 2 — Upload, overwrite, see versions
Upload a text file: test.txt with content "Version 1"Upload the same filename again: test.txt with content "Version 2" In the bucket → Show versions toggle (top right)You should see two versions of test.txt with different version IDsStep 3 — Delete and restore
Delete test.txt (without specifying a version ID)S3 creates a delete marker — the file appears gone Show versions → you will see the delete markerDelete the delete marker → file reappears This is how versioning protects against accidental deletes.Step 4 — Create a lifecycle rule
Bucket → Management → Create lifecycle ruleName: archive-logsPrefix: logs/ (only applies to objects in logs/ prefix) Add transitions: After 30 days → Standard-IA After 90 days → Glacier Flexible Retrieval Add expiration: After 365 days → expire current versions Scroll down → Add incomplete multipart upload action: After 7 days → delete incomplete uploads → Create ruleStep 5 — Generate a pre-signed URL
Using AWS CLI:aws s3 presign s3://devops-lab-2024-yourname/test.txt \ --expires-in 600 \ --region ap-south-1 Open the output URL in a browser — the file downloads without any login.Wait 10 minutes and try the URL again — it will show AccessDenied.Step 6 — Cleanup
Before deleting the bucket you must empty it:Bucket → Empty → confirm Then:Bucket → Delete → confirm bucket name → DeleteCommon Mistakes to Avoid
Common MistakeSetting the S3 access log destination to the same bucket being logged. Every access creates a log file, which is another access, which creates another log file. The bucket grows uncontrollably. Bills spike. Always use a separate dedicated logging bucket.
Common MistakeTurning off Block Public Access without realising what it does. Block Public Access protects the entire bucket even if a policy tries to make it public. Turning it off does not make the bucket public — but it means a bucket policy CAN make it public. Many engineers turn it off "to debug" and forget to turn it back on.
Common MistakeUsing SSE-KMS for high-throughput S3 workloads without checking KMS quotas. At millions of S3 operations per hour, KMS Decrypt is called for every download. The default KMS quota is 5,500-30,000 requests per second depending on region. Hitting the quota causes S3 throttling even though S3 itself is fine. Use SSE-S3 for high-throughput workloads unless audit trail is specifically required.
SecurityNever disable Block Public Access on a bucket unless you have a specific reason that justifies it. The number of data breaches from accidentally public S3 buckets has been enormous — enough that AWS made Block Public Access the default and added account-level enforcement. Treat any bucket with Block Public Access disabled as a security exception that needs review.
TipRun S3 Storage Lens on your account at least once. It shows you across all your buckets: which ones have versioning disabled, which have unencrypted objects, which have incomplete multipart uploads silently costing you money, and which storage classes your data is in. Most teams are surprised by what it finds — especially orphaned multipart uploads.