What you will learn
- Why AWS has many databases instead of one and the purpose-built philosophy
- RDS and Aurora — when you need relational with SQL joins and ACID transactions
- DynamoDB — when you need massive scale with simple access patterns
- DocumentDB — when your data is JSON documents with flexible schema
- Amazon Neptune — graph databases and relationship traversal
- Amazon Keyspaces — managed Cassandra for wide-column data
- Amazon Timestream — time-series data from IoT and metrics
- Amazon Redshift — data warehousing and analytical queries on large datasets
- ElastiCache — in-memory caching, not a primary database
- Amazon OpenSearch — full-text search and log analytics
- A simple decision framework to pick the right database every time
Why AWS Has So Many Databases
In the past, most applications used one relational database for everything. Orders went in. Products went in. User sessions went in. Search indexes went in. It worked — but it was never efficient. A relational database trying to do full-text search, graph traversal, time-series analytics, and high-speed caching is good at none of them.
AWS's philosophy is purpose-built databases — each one optimised for a specific data model and access pattern. Use the right tool for the job.
Relational data with joins → RDS or AuroraSimple key-value access at massive scale → DynamoDBJSON documents with flexible schema → DocumentDBRelationships between entities → NeptuneFull-text search → OpenSearchTime-series metrics → TimestreamAnalytical queries on billions of rows → RedshiftSub-millisecond caching → ElastiCacheThe goal: pick the database that fits the access pattern, not the database you already know.
Amazon RDS and Aurora — Relational
Use when:
Data has relationships that require SQL joinsACID transactions across multiple tables are requiredYour team knows SQL and relational modellingYou are migrating an existing relational application to AWSRDS engines: PostgreSQL, MySQL, MariaDB, Oracle, SQL Server Aurora: AWS-native PostgreSQL and MySQL compatible, 5x faster, 6 copies across 3 AZs
Real examples:
Zerodha trading ledger → RDS PostgreSQLDebits, credits, and balances across users, accounts, and tradesRelational integrity is non-negotiable — a sell without a corresponding credit would be a compliance violation Razorpay payment processing → Aurora MySQLMerchants, transactions, settlements, and disputes all linkedTransactions that span multiple tables require full ACID complianceWhen NOT to use:
Simple key lookups at millions of requests per second → DynamoDB is betterFlexible document data where schema changes frequently → DocumentDB is betterFull-text search on large text → OpenSearch is betterAmazon DynamoDB — NoSQL Key-Value
Use when:
Access pattern is simple — get item by ID, query by keyScale is massive — millions of reads/writes per secondLatency must be single-digit milliseconds at any scaleSchema evolves frequently or items have different attributesNeed global multi-region active-active writesReal examples:
Hotstar user sessions → DynamoDB100 million users, each session fetched by SessionIdSimple key lookup at massive scale — perfect DynamoDB fit CRED user reward points → DynamoDBUserId → current points balance + transaction historySimple key access, schema evolves as new reward types are added PhonePe UPI transaction state → DynamoDB + TTLTemporary transaction state expires automatically in 10 minutes via TTLWhen NOT to use:
Complex SQL queries with multiple joins → RDS is betterAd-hoc reporting on arbitrary attributes → not DynamoDB's strengthStrong relational integrity requirements → RDS is betterAmazon DocumentDB — JSON Documents
Use when:
Data is naturally JSON with flexible, nested structureYou are migrating a MongoDB application to AWSSchema needs to vary per document in the same collectionQueries are mostly within a single document, not across joinsDocumentDB is MongoDB-compatible. Your existing MongoDB queries, drivers, and tools work with minimal changes.
Real examples:
E-commerce product catalogue → DocumentDBEach product has different attributes:T-shirt: {size, color, material, fit}Phone: {storage, RAM, camera, battery, OS}No fixed schema needed — each document has its own structure User activity log → DocumentDBEach activity event has different fieldsFlexible schema evolves as the app adds new event typesWhen NOT to use:
Simple key-value access at massive scale → DynamoDB is faster and cheaperRelational data with joins → RDS is more appropriateTime-series data → Timestream is purpose-builtAmazon Neptune — Graph Database
Use when:
Data is about relationships between entitiesQueries traverse relationships across many nodesFinding patterns and connections is the primary use caseWhat makes graphs different:
Relational database: "find all friends of Rahul" → complex JOIN across tablesGraph database: "find all friends of Rahul" → one traversal, designed for this Relational: performance degrades as relationship depth increasesGraph: performance stays constant regardless of depthReal examples:
Social network → NeptuneRahul → follows → Priya → follows → Arjun"Find all people Rahul follows who also follow Arjun" → one graph query Fraud detection → Neptune"Find any accounts connected within 3 degrees to this known fraudulent account"Graph traversal finds these relationships instantly Recommendation engine (alternative) → Neptune"Users who bought X also bought Y" — find connections between products and usersSupported query languages: Gremlin and openCypher (SPARQL for RDF data)
Amazon Keyspaces — Wide-Column (Cassandra)
Use when:
You are already using Apache Cassandra on-premisesYou need massive write throughput with flexible columnsData is time-ordered and queries are by partition keyYou want Cassandra without managing clustersKeyspaces is managed Cassandra. Your existing CQL queries and drivers work without changes. AWS handles the infrastructure.
Difference from DynamoDB:
DynamoDB: AWS-proprietary, AWS-onlyKeyspaces: Cassandra-compatible, open standardChoose Keyspaces when migrating an existing Cassandra workload. Choose DynamoDB for new AWS-native NoSQL workloads.
Amazon Timestream — Time-Series Data
Use when:
Data is a sequence of measurements over timeQueries are mostly about time ranges and aggregations over timeSources are IoT sensors, application metrics, monitoring dataWhat makes time-series special:
Regular databases store every row equallyTime-series databases know that recent data is hot and old data is cold Timestream automatically: Keeps recent data in memory (fast access) Moves older data to magnetic storage (cheap) Deletes data beyond retention period automaticallyReal examples:
IoT temperature sensors → Timestream10,000 sensors sending readings every 30 secondsQuery: "average temperature in Mumbai zone B last 24 hours" Application performance metrics → TimestreamCPU, memory, latency metrics from hundreds of serversQuery: "p99 latency for orders API last 7 days" Financial tick data → TimestreamStock prices every secondQuery: "NIFTY50 closing price each day last year"100x faster and 10x cheaper than storing time-series data in a relational database for the same queries.
Amazon Redshift — Data Warehouse
Use when:
Analytical queries (OLAP) on very large datasets (billions of rows)Business intelligence and reportingComplex queries across many tables with aggregationsData arrives in batches, not real-time insertsOLTP vs OLAP:
OLTP (RDS, DynamoDB): Many small transactions (INSERT, UPDATE, DELETE) Few rows per query Real-time, transactional Application databases OLAP (Redshift): Few very large queries (GROUP BY, SUM, AVG across billions of rows) Millions of rows scanned per query Batch loaded, analytical Data warehouse for reportingRedshift Spectrum:
Query data stored in S3 directly from Redshift without loading it into Redshift first. Your S3 data lake becomes queryable with SQL.
SELECT * FROM spectrum.orders WHERE date > '2024-01-01'Data stays in S3. Redshift reads it directly. No ETL needed.ElastiCache — In-Memory Cache
Use when:
You need sub-millisecond response for frequently accessed dataYou want to reduce load on your primary databaseSession storage for stateless application serversElastiCache is not a primary database. It is a caching layer in front of your real database. Data in ElastiCache is temporary — it exists to speed up reads, not as a permanent record.
Redis vs Memcached:
Always choose Redis. It supports persistence, replication, Multi-AZ, pub/sub, and rich data types. Memcached is simpler but lacks all of these.
Amazon OpenSearch — Full-Text Search and Analytics
Use when:
Full-text search across large amounts of text dataLog analytics — querying application and infrastructure logsReal-time dashboards on streaming dataSearch with typo tolerance, relevance ranking, faceted filteringWhy not just use LIKE in SQL?
SELECT * FROM products WHERE name LIKE '%biryani%'Works on 1,000 rows. Terrible on 10 million rows.OpenSearch is built for search. Every field is indexed. Queries return in milliseconds regardless of data size. Results are ranked by relevance.
Real examples:
E-commerce search → OpenSearchUser types "biryani near me" with typos → OpenSearch returns relevant results Log analytics → OpenSearchShip all application logs → query in Kibana dashboard"Show all 5xx errors in the orders service last 2 hours" The ELK Stack: Elasticsearch (OpenSearch) + Logstash/Fluentd + KibanaThe Decision Framework
Answer these questions in order:
1. Do you need full-text search or log analytics? Yes → OpenSearch 2. Is this time-series data (metrics, IoT, monitoring)? Yes → Timestream 3. Is this analytical queries on billions of rows (BI, data warehouse)? Yes → Redshift 4. Do you need sub-millisecond caching in front of another database? Yes → ElastiCache 5. Is this graph data where relationships are the primary query pattern? Yes → Neptune 6. Are you migrating from MongoDB? Yes → DocumentDB 7. Are you migrating from Cassandra? Yes → Keyspaces 8. Do you need simple key-value access at massive scale (millions of req/sec)? Yes → DynamoDB 9. Do you need SQL, joins, ACID transactions, relational integrity? Yes → RDS (or Aurora for better performance)If you reach step 9 without a match, you almost certainly need RDS or Aurora.
Quick reference table:
| Database | Type | Best for |
|---|---|---|
| RDS / Aurora | Relational | SQL, joins, ACID, financial systems |
| DynamoDB | NoSQL Key-Value | Massive scale, simple key access, sessions |
| DocumentDB | NoSQL Document | JSON documents, MongoDB migration |
| Neptune | Graph | Social networks, fraud detection, relationships |
| Keyspaces | Wide-Column | Cassandra migration, wide-column IoT data |
| Timestream | Time-Series | IoT metrics, monitoring, stock prices |
| Redshift | Data Warehouse | Analytics, BI, reporting on billions of rows |
| ElastiCache | In-Memory Cache | Sub-ms reads, session store, DB offload |
| OpenSearch | Search and Analytics | Full-text search, log analytics, dashboards |
Hands-on Lab — Explore Each Service in the Console
This lab is a console tour — no resources created, no cost.
Step 1 — See the database landscape
AWS Console → type "database" in the search barNotice: RDS, DynamoDB, DocumentDB, Neptune, Keyspaces, Timestream, ElastiCache, Redshift, OpenSearch all appearEach is a separate service with its own consoleStep 2 — Compare RDS and DynamoDB pricing models
RDS → Create database (do not actually create)See: you must choose instance type and storage size upfrontThis is provisioned — you pay for what you allocate DynamoDB → Create table (do not actually create)See: On-Demand option — no capacity planning neededThis is serverless — you pay for what you useStep 3 — Open the Database Migration Service
DMS → Create replication instance (do not create)Notice: source and target endpoints can be different database typesOracle → Aurora: DMS handles the migration with SCT for schema conversionStep 4 — View Redshift Serverless
Redshift → Serverless dashboardSee: no cluster to provision — just query and pay per second of computeCompare with provisioned: provisioned requires choosing node type and countStep 5 — Use the decision framework
Think of a real application you are building or have built.Work through the decision framework questions above.Write down: which database does the framework point to? Do you agree? Why?No cleanup needed — this was a console tour, nothing was created.
Common Mistakes to Avoid
Common MistakeUsing RDS for everything out of familiarity. Relational databases are excellent for relational data. They are inefficient for simple key-value lookups at scale (DynamoDB is better), full-text search (OpenSearch is better), time-series metrics (Timestream is better), and analytical queries (Redshift is better). Using the right database cuts cost and improves performance.
Common MistakeUsing DynamoDB for data that has complex relationships. DynamoDB is optimised for simple key-based access. If your application needs to frequently join data across different entity types — users to orders to products to reviews — the access pattern fits RDS better. DynamoDB single-table design can handle some relational patterns but at the cost of significant complexity.
TipMost production applications use more than one database. Swiggy uses Aurora for the order transaction system (relational, ACID), DynamoDB for the restaurant catalogue (simple key access, schema flexibility), ElastiCache for active sessions (sub-millisecond), and OpenSearch for restaurant search (full-text). Each database handles what it is best at. The art is knowing which one to use for which part.