Skip to main content

AWS Database Selection Guide - Choosing the Right Database

Choose the right AWS database for every use case — RDS, Aurora, DynamoDB, DocumentDB, Neptune, ElastiCache, Redshift, Timestream, and OpenSearch compared.

What you will learn

  • Why AWS has many databases instead of one and the purpose-built philosophy
  • RDS and Aurora — when you need relational with SQL joins and ACID transactions
  • DynamoDB — when you need massive scale with simple access patterns
  • DocumentDB — when your data is JSON documents with flexible schema
  • Amazon Neptune — graph databases and relationship traversal
  • Amazon Keyspaces — managed Cassandra for wide-column data
  • Amazon Timestream — time-series data from IoT and metrics
  • Amazon Redshift — data warehousing and analytical queries on large datasets
  • ElastiCache — in-memory caching, not a primary database
  • Amazon OpenSearch — full-text search and log analytics
  • A simple decision framework to pick the right database every time

Why AWS Has So Many Databases

In the past, most applications used one relational database for everything. Orders went in. Products went in. User sessions went in. Search indexes went in. It worked — but it was never efficient. A relational database trying to do full-text search, graph traversal, time-series analytics, and high-speed caching is good at none of them.

AWS's philosophy is purpose-built databases — each one optimised for a specific data model and access pattern. Use the right tool for the job.

◈ DIAGRAM
Relational data with joins → RDS or Aurora
Simple key-value access at massive scale → DynamoDB
JSON documents with flexible schema → DocumentDB
Relationships between entities → Neptune
Full-text search → OpenSearch
Time-series metrics → Timestream
Analytical queries on billions of rows → Redshift
Sub-millisecond caching → ElastiCache

The goal: pick the database that fits the access pattern, not the database you already know.

Amazon RDS and Aurora — Relational

Use when:

TEXT
Data has relationships that require SQL joins
ACID transactions across multiple tables are required
Your team knows SQL and relational modelling
You are migrating an existing relational application to AWS

RDS engines: PostgreSQL, MySQL, MariaDB, Oracle, SQL Server Aurora: AWS-native PostgreSQL and MySQL compatible, 5x faster, 6 copies across 3 AZs

Real examples:

◈ DIAGRAM
Zerodha trading ledger → RDS PostgreSQL
Debits, credits, and balances across users, accounts, and trades
Relational integrity is non-negotiable — a sell without a corresponding credit would be a compliance violation
Razorpay payment processing → Aurora MySQL
Merchants, transactions, settlements, and disputes all linked
Transactions that span multiple tables require full ACID compliance

When NOT to use:

◈ DIAGRAM
Simple key lookups at millions of requests per second → DynamoDB is better
Flexible document data where schema changes frequently → DocumentDB is better
Full-text search on large text → OpenSearch is better

Amazon DynamoDB — NoSQL Key-Value

Use when:

TEXT
Access pattern is simple — get item by ID, query by key
Scale is massive — millions of reads/writes per second
Latency must be single-digit milliseconds at any scale
Schema evolves frequently or items have different attributes
Need global multi-region active-active writes

Real examples:

◈ DIAGRAM
Hotstar user sessions → DynamoDB
100 million users, each session fetched by SessionId
Simple key lookup at massive scale — perfect DynamoDB fit
CRED user reward points → DynamoDB
UserId → current points balance + transaction history
Simple key access, schema evolves as new reward types are added
PhonePe UPI transaction state → DynamoDB + TTL
Temporary transaction state expires automatically in 10 minutes via TTL

When NOT to use:

◈ DIAGRAM
Complex SQL queries with multiple joins → RDS is better
Ad-hoc reporting on arbitrary attributes → not DynamoDB's strength
Strong relational integrity requirements → RDS is better

Amazon DocumentDB — JSON Documents

Use when:

TEXT
Data is naturally JSON with flexible, nested structure
You are migrating a MongoDB application to AWS
Schema needs to vary per document in the same collection
Queries are mostly within a single document, not across joins

DocumentDB is MongoDB-compatible. Your existing MongoDB queries, drivers, and tools work with minimal changes.

Real examples:

◈ DIAGRAM
E-commerce product catalogue → DocumentDB
Each product has different attributes:
T-shirt: {size, color, material, fit}
Phone: {storage, RAM, camera, battery, OS}
No fixed schema needed — each document has its own structure
User activity log → DocumentDB
Each activity event has different fields
Flexible schema evolves as the app adds new event types

When NOT to use:

◈ DIAGRAM
Simple key-value access at massive scale → DynamoDB is faster and cheaper
Relational data with joins → RDS is more appropriate
Time-series data → Timestream is purpose-built

Amazon Neptune — Graph Database

Use when:

TEXT
Data is about relationships between entities
Queries traverse relationships across many nodes
Finding patterns and connections is the primary use case

What makes graphs different:

◈ DIAGRAM
Relational database: "find all friends of Rahul" → complex JOIN across tables
Graph database: "find all friends of Rahul" → one traversal, designed for this
Relational: performance degrades as relationship depth increases
Graph: performance stays constant regardless of depth

Real examples:

◈ DIAGRAM
Social network → Neptune
Rahul → follows → Priya → follows → Arjun
"Find all people Rahul follows who also follow Arjun" → one graph query
Fraud detection → Neptune
"Find any accounts connected within 3 degrees to this known fraudulent account"
Graph traversal finds these relationships instantly
Recommendation engine (alternative) → Neptune
"Users who bought X also bought Y" — find connections between products and users

Supported query languages: Gremlin and openCypher (SPARQL for RDF data)

Amazon Keyspaces — Wide-Column (Cassandra)

Use when:

TEXT
You are already using Apache Cassandra on-premises
You need massive write throughput with flexible columns
Data is time-ordered and queries are by partition key
You want Cassandra without managing clusters

Keyspaces is managed Cassandra. Your existing CQL queries and drivers work without changes. AWS handles the infrastructure.

Difference from DynamoDB:

TEXT
DynamoDB: AWS-proprietary, AWS-only
Keyspaces: Cassandra-compatible, open standard

Choose Keyspaces when migrating an existing Cassandra workload. Choose DynamoDB for new AWS-native NoSQL workloads.

Amazon Timestream — Time-Series Data

Use when:

TEXT
Data is a sequence of measurements over time
Queries are mostly about time ranges and aggregations over time
Sources are IoT sensors, application metrics, monitoring data

What makes time-series special:

TEXT
Regular databases store every row equally
Time-series databases know that recent data is hot and old data is cold
Timestream automatically:
Keeps recent data in memory (fast access)
Moves older data to magnetic storage (cheap)
Deletes data beyond retention period automatically

Real examples:

◈ DIAGRAM
IoT temperature sensors → Timestream
10,000 sensors sending readings every 30 seconds
Query: "average temperature in Mumbai zone B last 24 hours"
Application performance metrics → Timestream
CPU, memory, latency metrics from hundreds of servers
Query: "p99 latency for orders API last 7 days"
Financial tick data → Timestream
Stock prices every second
Query: "NIFTY50 closing price each day last year"

100x faster and 10x cheaper than storing time-series data in a relational database for the same queries.

Amazon Redshift — Data Warehouse

Use when:

TEXT
Analytical queries (OLAP) on very large datasets (billions of rows)
Business intelligence and reporting
Complex queries across many tables with aggregations
Data arrives in batches, not real-time inserts

OLTP vs OLAP:

TEXT
OLTP (RDS, DynamoDB):
Many small transactions (INSERT, UPDATE, DELETE)
Few rows per query
Real-time, transactional
Application databases
OLAP (Redshift):
Few very large queries (GROUP BY, SUM, AVG across billions of rows)
Millions of rows scanned per query
Batch loaded, analytical
Data warehouse for reporting

Redshift Spectrum:

Query data stored in S3 directly from Redshift without loading it into Redshift first. Your S3 data lake becomes queryable with SQL.

TEXT
SELECT * FROM spectrum.orders WHERE date > '2024-01-01'
Data stays in S3. Redshift reads it directly. No ETL needed.

ElastiCache — In-Memory Cache

Use when:

TEXT
You need sub-millisecond response for frequently accessed data
You want to reduce load on your primary database
Session storage for stateless application servers

ElastiCache is not a primary database. It is a caching layer in front of your real database. Data in ElastiCache is temporary — it exists to speed up reads, not as a permanent record.

Redis vs Memcached:

Always choose Redis. It supports persistence, replication, Multi-AZ, pub/sub, and rich data types. Memcached is simpler but lacks all of these.

Amazon OpenSearch — Full-Text Search and Analytics

Use when:

TEXT
Full-text search across large amounts of text data
Log analytics — querying application and infrastructure logs
Real-time dashboards on streaming data
Search with typo tolerance, relevance ranking, faceted filtering

Why not just use LIKE in SQL?

TEXT
SELECT * FROM products WHERE name LIKE '%biryani%'
Works on 1,000 rows. Terrible on 10 million rows.

OpenSearch is built for search. Every field is indexed. Queries return in milliseconds regardless of data size. Results are ranked by relevance.

Real examples:

◈ DIAGRAM
E-commerce search → OpenSearch
User types "biryani near me" with typos → OpenSearch returns relevant results
Log analytics → OpenSearch
Ship all application logs → query in Kibana dashboard
"Show all 5xx errors in the orders service last 2 hours"
The ELK Stack: Elasticsearch (OpenSearch) + Logstash/Fluentd + Kibana

The Decision Framework

Answer these questions in order:

◈ DIAGRAM
1. Do you need full-text search or log analytics?
Yes → OpenSearch
2. Is this time-series data (metrics, IoT, monitoring)?
Yes → Timestream
3. Is this analytical queries on billions of rows (BI, data warehouse)?
Yes → Redshift
4. Do you need sub-millisecond caching in front of another database?
Yes → ElastiCache
5. Is this graph data where relationships are the primary query pattern?
Yes → Neptune
6. Are you migrating from MongoDB?
Yes → DocumentDB
7. Are you migrating from Cassandra?
Yes → Keyspaces
8. Do you need simple key-value access at massive scale (millions of req/sec)?
Yes → DynamoDB
9. Do you need SQL, joins, ACID transactions, relational integrity?
Yes → RDS (or Aurora for better performance)

If you reach step 9 without a match, you almost certainly need RDS or Aurora.

Quick reference table:

Database Type Best for
RDS / Aurora Relational SQL, joins, ACID, financial systems
DynamoDB NoSQL Key-Value Massive scale, simple key access, sessions
DocumentDB NoSQL Document JSON documents, MongoDB migration
Neptune Graph Social networks, fraud detection, relationships
Keyspaces Wide-Column Cassandra migration, wide-column IoT data
Timestream Time-Series IoT metrics, monitoring, stock prices
Redshift Data Warehouse Analytics, BI, reporting on billions of rows
ElastiCache In-Memory Cache Sub-ms reads, session store, DB offload
OpenSearch Search and Analytics Full-text search, log analytics, dashboards

Hands-on Lab — Explore Each Service in the Console

This lab is a console tour — no resources created, no cost.

Step 1 — See the database landscape

◈ DIAGRAM
AWS Console → type "database" in the search bar
Notice: RDS, DynamoDB, DocumentDB, Neptune, Keyspaces, Timestream, ElastiCache, Redshift, OpenSearch all appear
Each is a separate service with its own console

Step 2 — Compare RDS and DynamoDB pricing models

◈ DIAGRAM
RDS → Create database (do not actually create)
See: you must choose instance type and storage size upfront
This is provisioned — you pay for what you allocate
DynamoDB → Create table (do not actually create)
See: On-Demand option — no capacity planning needed
This is serverless — you pay for what you use

Step 3 — Open the Database Migration Service

◈ DIAGRAM
DMS → Create replication instance (do not create)
Notice: source and target endpoints can be different database types
Oracle → Aurora: DMS handles the migration with SCT for schema conversion

Step 4 — View Redshift Serverless

◈ DIAGRAM
Redshift → Serverless dashboard
See: no cluster to provision — just query and pay per second of compute
Compare with provisioned: provisioned requires choosing node type and count

Step 5 — Use the decision framework

TEXT
Think of a real application you are building or have built.
Work through the decision framework questions above.
Write down: which database does the framework point to? Do you agree? Why?

No cleanup needed — this was a console tour, nothing was created.

Common Mistakes to Avoid

Common Mistake

Using RDS for everything out of familiarity. Relational databases are excellent for relational data. They are inefficient for simple key-value lookups at scale (DynamoDB is better), full-text search (OpenSearch is better), time-series metrics (Timestream is better), and analytical queries (Redshift is better). Using the right database cuts cost and improves performance.

Common Mistake

Using DynamoDB for data that has complex relationships. DynamoDB is optimised for simple key-based access. If your application needs to frequently join data across different entity types — users to orders to products to reviews — the access pattern fits RDS better. DynamoDB single-table design can handle some relational patterns but at the cost of significant complexity.

Tip

Most production applications use more than one database. Swiggy uses Aurora for the order transaction system (relational, ACID), DynamoDB for the restaurant catalogue (simple key access, schema flexibility), ElastiCache for active sessions (sub-millisecond), and OpenSearch for restaurant search (full-text). Each database handles what it is best at. The art is knowing which one to use for which part.

Resources

AWS Direct Connect vs Site-to-Site VPN Failover

AWS Direct Connect vs Site-to-Site VPN Failover

Direct Connect vs VPN isn't really either/or for production — it's a primary-plus-failover pattern. Here's how to design it, and when either/or is right.

5 min read•Aug 2026
Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot, at the crossover where Lambda stops being cheaper — 2026 pricing, invocation thresholds, and interruption math.

5 min read•Aug 2026
Secrets Manager vs Parameter Store vs Vault

Secrets Manager vs Parameter Store vs Vault

AWS Secrets Manager, Parameter Store, and HashiCorp Vault compared for 2026 - cost math, rotation, multi-cloud fit, and the Vault-to-OpenBao fork.

5 min read•Aug 2026
AWS VPC Security: Hardening Every Layer

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

5 min read•Jul 2026
Event-Driven Architecture on AWS Explained

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

5 min read•Jul 2026
S3 vs RDS vs DynamoDB: Choosing AWS Storage

S3 vs RDS vs DynamoDB: Choosing AWS Storage

Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.

5 min read•Jul 2026
AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

5 min read•Jul 2026
EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2, Lambda, or Fargate — choosing the wrong AWS compute option costs you money and performance. Here is exactly when to use each one in production.

5 min read•Jul 2026

Explore More in AWS Storage and Databases

All 6 Topics

Frequently Asked Questions

Is AWS Database Selection Guide - Choosing the Right Database free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the AWS Database Selection Guide - Choosing the Right Database topic cover?

Choose the right AWS database for every use case — RDS, Aurora, DynamoDB, DocumentDB, Neptune, ElastiCache, Redshift, Timestream, and OpenSearch compared.