S3 vs RDS vs DynamoDB: Choosing AWS Storage
Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.
A team at a Bengaluru e-commerce startup stored their product catalogue in RDS. 2 million products, each with different attributes — t-shirts have size and colour, phones have RAM and storage, books have ISBN and author. Every attribute was a column. Their schema had 47 columns and half of them were NULL for any given product. Queries were slow. Adding a new product type required a schema migration. They migrated to DynamoDB and query times dropped from 180ms to 8ms.
The same team stored user profile images in DynamoDB. Images are binary objects. DynamoDB items have a 400 KB limit per item. They stored base64-encoded thumbnails in DynamoDB items. It worked until they added high-resolution images. They migrated to S3 and API response times dropped from 1.2 seconds to 90ms, because S3 returns direct presigned URLs rather than binary data through the API.
The right storage choice is never about which service is "better." It is always about which service matches your data model and access pattern. I should flag upfront that the specific numbers in this article (latency figures, storage costs, illustrative company examples) are meant to be directionally correct and educational rather than precise current benchmarks — you should verify exact current pricing directly on AWS's pricing pages before budgeting, since AWS revises rates periodically.
The Three Storage Models
Each service is built around a fundamentally different model:
S3 -- Object Storage:Any file, any size, stored as a named object in a bucketAccess via HTTP API: GET /bucket/key -> download the fileNo schema. No relationships. No queries on file contents.Best for: files, media, backups, logs, large datasets RDS -- Relational Database:Tables with fixed columns, rows with valuesAccess via SQL: SELECT, JOIN, GROUP BY, WHERESchema enforced. ACID transactions. Complex queries.Best for: structured data with relationships and complex queries DynamoDB -- Key-Value / Document:Tables with items, flexible attributes per itemAccess via key: get item by partition key + optional sort keyNo schema enforcement. Millisecond latency at any scale.Best for: simple key-based access at massive or unpredictable scaleThe mistake is trying to use one service for all three purposes.
When S3 is the Right Choice
S3 is for files and objects — anything you store as a complete unit and retrieve as a complete unit.
At a large streaming platform like Hotstar, video content at petabyte scale lives in S3. Videos are retrieved as complete files for streaming, not queried by attribute. S3 stores them durably across multiple Availability Zones and delivers them through CloudFront to edge locations worldwide. No database could store or serve video content at this cost or scale.
S3 is correct for:
- Media files — images, videos, audio, documents
- Application logs — CloudTrail logs, access logs, application output
- Data lake raw storage — CSV, JSON, Parquet files for Athena to query
- Static website assets — HTML, CSS, JavaScript
- Backup snapshots — RDS snapshots exported to S3, EBS snapshots
- Large datasets for batch processing — EMR and Glue read from S3
S3 is wrong when:
- You need to query the contents of files with WHERE clauses — use RDS or DynamoDB
- You need to update part of an object — S3 objects are replaced, not updated
- You need sub-millisecond latency on frequent small reads — use ElastiCache or DynamoDB
- You need relational integrity — S3 has no concept of foreign keys or transactions
S3 cost reality (verify current rates on AWS's pricing page before budgeting):
S3 Standard: roughly $0.023/GB/month in most US regionsS3 Standard-IA: roughly $0.0125/GB/month, plus a per-GB retrieval feeS3 Glacier Flexible: a fraction of Standard, plus retrieval fees and a wait time Set lifecycle rules from day one. Never pay Standard prices for data nobody reads.When RDS is the Right Choice
RDS is for structured relational data where your queries cannot be predicted at schema design time.
A trading platform's ledger has debits, credits, balances, accounts, instruments, and orders. Each trade creates entries across multiple tables atomically. A credit to one account is simultaneously a debit to another. This is a multi-table transaction that must either fully succeed or fully fail. DynamoDB's transaction support exists but is more limited and comes at extra cost per transaction. PostgreSQL on RDS handles this with native ACID transactions, familiar SQL, and the ability to run arbitrary queries against the data that the engineering team has not thought of yet.
RDS is correct for:
- Financial records requiring ACID transactions across multiple tables
- Reporting systems that run ad-hoc GROUP BY and JOIN queries
- User management systems with complex relational structures
- Any domain where the access patterns are not fully known at design time
- Applications migrating from on-premises MySQL or PostgreSQL
RDS is wrong when:
- You need very large storage scale or automatic horizontal scaling — Aurora or DynamoDB scale further
- You need automatic scaling to millions of requests per second — DynamoDB handles this
- Schema changes happen frequently — RDS schema migrations require careful planning and downtime management
- Your access pattern is always "get item by ID" — DynamoDB is typically cheaper and faster for this narrow pattern
RDS vs Aurora: Aurora is generally the better default for new production workloads on AWS — it's built for higher throughput and better multi-AZ durability than standard RDS MySQL/PostgreSQL, at a modest cost premium. Verify current instance pricing for your specific class before committing, since Aurora pricing has its own tiering separate from standard RDS.
RDS connection management: RDS has a maximum connection limit based on instance size. At scale, Lambda functions and containerised services can exhaust this limit. Use RDS Proxy — it pools connections and multiplexes many concurrent Lambda invocations through a smaller pool of real database connections.
When DynamoDB is the Right Choice
DynamoDB is for simple, fast, key-based access at any scale — from one request per day to millions per second.
At a UPI-scale payments platform, every transaction creates a state record — a temporary object with status (initiated, processing, completed, failed) that must be checked in single-digit milliseconds. Millions of transactions per day. Each state record is accessed by transaction ID only — no joins, no aggregations. DynamoDB returns any item by primary key in low single-digit milliseconds consistently, regardless of whether the table has 1,000 items or 10 billion items. No relational database matches this at scale without significant added architecture complexity.
DynamoDB is correct for:
- Session management — user sessions accessed by session token, auto-expire via TTL
- Leaderboards — sorted sets by score using Sort Key
- Shopping carts — access by user ID, frequent reads and writes
- IoT device state — access by device ID, frequent small updates
- High-volume event logging — write millions of events per second
- Real-time inventory — global multi-region active-active writes with Global Tables
DynamoDB is wrong when:
- You need SQL GROUP BY across all items — Scan operations are slow and expensive
- You need complex JOIN queries across different entity types — RDS is the right model
- Individual item size exceeds 400 KB — use S3 for the data, DynamoDB for the metadata
- Access patterns change frequently — DynamoDB schema design is access-pattern-driven, and redesigning it later is painful
DynamoDB cost: AWS cut on-demand throughput pricing roughly in half in November 2024, and current on-demand rates (verify against AWS's live pricing page, since these are the kind of numbers that go stale fast) are approximately $0.125 per million read request units and $0.625 per million write request units for Standard tables, plus roughly $0.25/GB-month for storage. A useful mental model from current cost-analysis writeups: at 100 million reads and 20 million writes per day, on-demand mode runs meaningfully higher than provisioned capacity with auto-scaling once your traffic is predictable enough to forecast — the commonly cited savings from switching is in the 60-80% range, though you should model your own traffic rather than assume that ratio applies to you. Each Global Secondary Index you add roughly doubles the write cost for that index, since every write also updates the index.
Comparing the Three Services Directly
Data Shape and Access:
| Dimension | S3 | RDS |
|---|---|---|
| Data type | Files and objects | Structured rows and columns |
| Query language | None (key-based GET) | SQL — full JOIN, GROUP BY |
| Max item/row size | 5 TB per object | Limited by disk |
| Latency | Tens to low hundreds of ms | Single-digit to tens of ms (indexed query) |
| Scaling model | Unlimited, automatic | Vertical + Read Replicas |
| ACID transactions | No | Yes |
| Schema | None | Fixed — migrations required |
Data Shape and Access, continued:
| Dimension | DynamoDB |
|---|---|
| Data type | Items with flexible attributes |
| Query language | Key-based — partition key + sort key |
| Max item/row size | 400 KB per item |
| Latency | Single-digit ms (key lookup) |
| Scaling model | Automatic, unlimited |
| ACID transactions | Yes (limited, extra cost) |
| Schema | Flexible per item |
Combining All Three in One Architecture
The most effective architectures use all three together, each handling what it does best. A food-delivery platform at Swiggy's scale is a representative pattern:
S3: Restaurant menu photos and videos Invoice PDFs for completed orders Application logs for analytics RDS (Aurora MySQL): Orders table (order_id, user_id, restaurant_id, total, status) Restaurants table (restaurant_id, name, location, rating) Complex queries: GROUP BY restaurant for analytics, JOIN for order history DynamoDB: Active session tokens (expire via TTL after 24 hours) Real-time delivery driver locations (updated every few seconds per driver) Rate limiting counters (user X made N API calls in the last minute)None of these are interchangeable. Storing restaurant photos in RDS or DynamoDB would be architecturally wrong. Storing driver location in RDS would be too slow at that update frequency. Storing order history in DynamoDB would make reporting queries painful.
Trade-offs and the Decision You Make
Requirement Fit — S3 and RDS:
| Requirement | Use S3 | Use RDS |
|---|---|---|
| Store images, videos, files | Yes | No |
| SQL joins and GROUP BY | No | Yes |
| Get item by ID in single-digit ms | No | Maybe |
| ACID cross-table transactions | No | Yes |
| Store 100 TB+ at low cost | Yes | No |
| Unknown query patterns | No | Yes |
Requirement Fit — DynamoDB:
| Requirement | Use DynamoDB |
|---|---|
| Store images, videos, files | No |
| SQL joins and GROUP BY | No |
| Get item by ID in single-digit ms | Yes |
| ACID cross-table transactions | Partial |
| Store 100 TB+ at low cost | Yes |
| Unknown query patterns | No |
Production Implementation Guidelines
Follow these rules before finalising your storage choice:
- Identify every access pattern before choosing a database. If you cannot enumerate your access patterns, RDS is the safer default — it allows ad-hoc queries. DynamoDB requires knowing access patterns upfront.
- Store binary objects — images, videos, documents — in S3 and reference them by URL. Never store binary data in a relational database or DynamoDB if the objects exceed a few KB.
- Start with DynamoDB On-Demand for new, unpredictable workloads. Switch to Provisioned with auto-scaling once traffic is predictable — the savings are usually substantial, but model your own traffic before assuming a specific percentage.
- Use ElastiCache in front of RDS for read-heavy workloads. An RDS query returning the same product record thousands of times per minute is wasteful. Cache it in Redis for sub-millisecond retrieval.
- Never run analytics directly on production RDS. Use Read Replicas for reporting, or export data to Redshift for complex BI queries.
- Audit your Global Secondary Indexes on DynamoDB tables periodically — each one adds ongoing write cost, and it's easy to accumulate indexes that are no longer used by any live access pattern.
NoteReferences and Further Reading
- AWS database selection guide — AWS's own comparison of all database services
- DynamoDB best practices — single-table design and access pattern design
- Amazon DynamoDB Pricing — current on-demand and provisioned rates, since pricing changes over time
Frequently Asked Questions
Can you store large binary files like images directly in DynamoDB or RDS?
Not effectively — DynamoDB items have a 400 KB limit per item, and while RDS technically supports BLOB columns, storing large binary objects there is architecturally the wrong tool. S3 is built specifically for files and objects, returning direct presigned URLs instead of forcing binary data through an API response.
When is DynamoDB the wrong choice even though it's the fastest of the three for key lookups?
When you need SQL-style GROUP BY or JOIN queries across different entity types, or when your access patterns aren't fully known upfront — DynamoDB's Scan operations are slow and expensive for ad-hoc queries, and its schema design is access-pattern-driven, making changes painful later.
Why did switching a product catalogue from RDS to DynamoDB cut query times so dramatically in the example?
Because the data had wildly inconsistent attributes across product types (t-shirts need size/colour, phones need RAM/storage), which meant a sparse, mostly-NULL relational schema — DynamoDB's flexible per-item attributes are a much better structural fit for that kind of heterogeneous data.
Does adding a Global Secondary Index to a DynamoDB table have a meaningful cost impact?
Yes — each GSI roughly doubles the write cost for that index, since every write also updates the index. It's worth periodically auditing GSIs on production tables for ones no longer used by any live access pattern.
Should production reporting queries run directly against your primary RDS instance?
No — use Read Replicas for reporting, or export data to Redshift for complex BI queries. Running analytics directly on production RDS risks resource contention with the transactional workload the database is actually serving.
Discussion0