Skip to main content

Amazon DynamoDB - Serverless NoSQL at Scale

Design DynamoDB tables with the right primary key, capacity mode, and access patterns, and use Streams, Global Tables, DAX, and TTL for production workloads.

What you will learn

  • What DynamoDB is and when NoSQL is the right choice over relational databases
  • Tables, items, attributes — and how flexible schema actually works in practice
  • Primary key design — Partition Key only vs Partition Key + Sort Key
  • Why partition key selection determines the performance ceiling of your entire table
  • On-Demand vs Provisioned capacity modes and when each saves money
  • DAX — microsecond reads with zero application code changes
  • DynamoDB Streams — the ordered changelog that powers event-driven architecture
  • Global Tables — active-active multi-region replication and the Streams prerequisite
  • TTL — automatic item expiry with zero extra cost
  • PITR and on-demand backups for disaster recovery
  • S3 export for analytics without consuming read capacity

Why this matters

Hotstar serves metadata for millions of videos — episode details, thumbnails, descriptions. The access pattern is simple: get video by ID. The scale is massive: millions of reads per second with single-digit millisecond response time required. RDS cannot handle this without enormous cost and operational complexity. CRED stores user reward points and transaction history — flexible schema evolves as the product adds new reward types without schema migrations. PhonePe processes UPI transactions where session tokens must expire automatically after 10 minutes — TTL handles this at zero cost. DynamoDB is not a replacement for relational databases. It is the correct choice when the access pattern is simple, the scale is massive, and the schema evolves rapidly.

What is DynamoDB and When to Use It

DynamoDB is AWS's fully managed proprietary NoSQL database. Millisecond latency at any scale, serverless, Multi-AZ by default. No schema enforcement — each item can have different attributes.

TEXT
When DynamoDB wins:
Data does not need joins across multiple tables
Schema evolves frequently — agile development, new features weekly
Need single-digit millisecond reads at massive scale
Serverless — no capacity planning needed (On-Demand mode)
Need global multi-region active-active writes (Global Tables)
Session data, tokens, temporary records with automatic expiry (TTL)
When DynamoDB loses:
You need SQL joins across multiple entities
Complex reporting queries with GROUP BY, aggregations
ACID transactions across many tables (use RDS instead)
Ad-hoc queries on arbitrary fields you have not indexed

DynamoDB vs RDS — when to choose which:

Question If Yes → Use
Do your entities have complex relationships? RDS
Do you need SQL joins in production queries? RDS
Do you need complex aggregations and reporting? RDS
Is your scale massive with simple key-based access? DynamoDB
Does your schema change frequently? DynamoDB
Do you need global multi-region active-active writes? DynamoDB
Do you need millisecond latency at any scale serverlessly? DynamoDB

Tables, Items, Attributes, and Primary Keys

Key terms:

◈ DIAGRAM
Table → named container for related data, like a database table
No schema enforcement — only rule is every item must have the Primary Key
Item → one record, like a row in SQL
Maximum size per item: 400 KB
Attribute → a field of an item, like a column
Different items can have completely different attributes

Flexible attributes in practice:

◈ DIAGRAM
RDS — every row must have the same columns:
id name email age
1 Rahul rahul@devops.in 25
2 Priya priya@devops.in NULL ← age must exist even if null
DynamoDB — each item has only what it needs:
Item 1: { id: "1", name: "Rahul", email: "rahul@devops.in", age: 25 }
Item 2: { id: "2", name: "Priya", phone: "9876543210", city: "Pune" }
Item 3: { id: "3", name: "Arjun", score: 99, level: 5, badges: ["gold","silver"] }

Item 2 has phone and city but no email. Item 3 has score, level, and a list of badges. All in the same table, no errors. This is why DynamoDB enables rapid schema evolution — new attributes just appear on new items.

Data types:

TEXT
Scalar: String, Number, Binary, Boolean, Null
Document: List, Map (nested structures like JSON objects)
Set: String Set, Number Set, Binary Set

Primary Key — decided at table creation, cannot be changed:

Two options:

TEXT
Option 1 — Partition Key only (simple primary key):
One attribute must be unique per item
Example: UserId — each user has one item
Option 2 — Partition Key + Sort Key (composite primary key):
Combination of two attributes must be unique
Example: UserId + GameId — same user can have many game scores

Game scores table example:

◈ DIAGRAM
Partition Key Sort Key Score Result
(UserId) (GameId)
user-001 game-42 92 Win
user-002 game-18 14 Lose
user-002 game-42 77 Win ← same user, different game

User-002 appears twice — same user played two games. The UserId + GameId combination is unique so both rows are valid. Sort Key allows one entity to have many related items in the same table.

Partition Key Design — The Most Critical Decision

The partition key determines which physical partition holds your data. A poor choice causes hot partitions — one partition gets all the traffic while others sit idle. This is the most common DynamoDB performance mistake.

◈ DIAGRAM
Good partition keys (high cardinality, even distribution):
UserId → millions of unique users, traffic spread evenly
OrderId → unique per order, completely random distribution
DeviceId → millions of IoT devices, random access pattern
SessionToken → random UUID, perfectly distributed
Bad partition keys (low cardinality, hot partitions):
Status → only a few values (active, inactive) — all active users on one partition
Country → only a few hundred values — India has far more traffic than others
Boolean flag → only two values — true gets all writes if most items are true
Date → today's date is all writes — yesterday's date gets nothing
Common Mistake

Using a low-cardinality value like Status or Country as the partition key. If 90% of your writes have Status=active, then 90% of your writes go to a single partition. DynamoDB throttles that partition. Your application sees errors. Fix by choosing a high-cardinality attribute like UserId or using partition key sharding.

Read and Write Capacity Modes

On-Demand Mode:

DynamoDB automatically scales reads and writes up and down instantly. No capacity planning needed.

TEXT
Pay only for what you actually use
Handles any traffic without throttling or configuration
More expensive per request than Provisioned
Best for: unpredictable traffic, new apps with unknown usage, spiky workloads

Provisioned Mode:

You specify upfront how many reads per second (RCU) and writes per second (WCU).

◈ DIAGRAM
RCU = Read Capacity Unit = one strongly consistent read per second for items up to 4 KB
WCU = Write Capacity Unit = one write per second for items up to 1 KB
Can add Auto Scaling to adjust RCU and WCU automatically within set limits
Cheaper per request than On-Demand for predictable steady workloads
Risk: under-provision → throttling; over-provision → wasted cost
Provisioned Mode On-Demand Mode
Capacity planning Required None
Cost per request Cheaper More expensive
Throttling risk Yes if under-provisioned No
Auto Scaling Optional add-on Built-in automatic
Best for Predictable steady traffic Unpredictable or spiky traffic
Tip

Start with On-Demand for new applications — no guessing required. Once you understand your traffic pattern after a few weeks, switch to Provisioned with Auto Scaling to save cost.

DAX — DynamoDB Accelerator

DynamoDB responds in single-digit milliseconds. For read-heavy applications hitting the same data constantly — a leaderboard, product catalog, popular items — those milliseconds add up and the database gets congested with repeated reads.

DAX is a fully managed in-memory cache specifically for DynamoDB. Frequently accessed data is returned in microseconds — up to 10x faster than going to DynamoDB directly.

◈ DIAGRAM
Without DAX:
Every request → DynamoDB → single-digit milliseconds
1000 requests for same product → 1000 DynamoDB reads
With DAX:
First request → DAX miss → DynamoDB → cached in DAX
Next 999 requests → DAX hit → microseconds, DynamoDB never touched

No application code changes needed:

DAX uses the same DynamoDB API. You point your SDK at the DAX cluster endpoint instead of the DynamoDB endpoint. Everything else stays the same.

TEXT
Old endpoint: dynamodb.ap-south-1.amazonaws.com
New endpoint: my-dax-cluster.abc123.dax-clusters.ap-south-1.amazonaws.com

DAX vs ElastiCache for DynamoDB:

TEXT
DAX:
Purpose-built for DynamoDB
Zero code changes — same API
Caches individual items and query results automatically
Use when: repeated reads of the same DynamoDB items
ElastiCache:
General purpose cache
Requires code changes — check cache first manually
Caches computed results, aggregations
Use when: need to cache complex computations or results DynamoDB cannot produce
Tip

For DynamoDB read caching, always try DAX first. It requires no code changes and gives microsecond reads. Only reach for ElastiCache if you need to cache something DAX cannot — like an aggregated result computed across many items.

DynamoDB Streams

Every time an item is created, updated, or deleted, DynamoDB Streams captures that change as an ordered event. Think of it as a real-time changelog of everything that happened in your table.

◈ DIAGRAM
Application creates new user in DynamoDB
↓
DynamoDB Streams captures the INSERT event
↓
Lambda triggered by the stream
↓
Lambda sends welcome email via SES
↓
No polling, no scheduled jobs — fully event-driven

Two ways to consume streams:

TEXT
DynamoDB Streams (original):
24-hour retention
Limited consumers
Process with Lambda or Kinesis Adapter
Good for simple real-time triggers
Kinesis Data Streams (newer, more powerful):
1-year retention
High number of consumers
Process with Lambda, Kinesis Analytics, Firehose, Glue Streaming
Good for complex pipelines, analytics, long-term archiving

Full stream processing pattern:

◈ DIAGRAM
Application writes to DynamoDB
↓
DynamoDB Streams
├── Lambda → send notification, update derived table
├── Kinesis Data Streams → Firehose → S3 (archive)
└── Kinesis Data Streams → Redshift (analytics)

Global Tables — Active-Active Multi-Region

Global Tables makes your DynamoDB table available in multiple regions simultaneously with active-active replication — writes accepted in any region, replicated everywhere.

◈ DIAGRAM
ap-south-1 (Mumbai) us-east-1 (Virginia)
User in India writes User in US writes
to Mumbai replica ←→ to Virginia replica
↓ ↓
Change replicates Change replicates
to Virginia to Mumbai
within seconds within seconds

Both regions are fully active. A user in Mumbai reads from and writes to the Mumbai replica with single-digit millisecond latency. A user in New York does the same against the Virginia replica.

DynamoDB Streams must be enabled before creating Global Tables:

Global Tables uses Streams internally as the replication engine — every change recorded in Streams is what gets replicated to other regions.

◈ DIAGRAM
Create table → Enable DynamoDB Streams → Create Global Table → Add regions
Remember

DynamoDB Streams is a hard prerequisite for Global Tables. If you try to create a Global Table without Streams enabled, it fails. Enable Streams first, always.

Global Tables vs RDS Cross-Region Read Replica:

DynamoDB Global Tables RDS Cross-Region Read Replica
Write regions All regions (active-active) One primary only (active-passive)
Replication lag Sub-second Minutes
Conflicts Last-writer-wins Not applicable
Use case Global apps needing local writes Read scaling in another region

TTL — Time To Live

TTL automatically deletes items from your table after a set time. No cleanup code, no scheduled jobs, no extra cost.

How it works:

◈ DIAGRAM
You choose any attribute name — ExpTime, TTL, ExpiresAt, anything
Tell DynamoDB once in table settings: this is my TTL attribute
When creating an item, set that attribute to an epoch timestamp
DynamoDB scans the attribute vs current time
When current time passes the timestamp → item automatically deleted
Free — does not consume write capacity

Session management example:

◈ DIAGRAM
User logs in → create session item with ExpTime = now + 86400 (24 hours)
User active → session item exists, valid
24 hours pass → DynamoDB automatically deletes the session item
No cleanup Lambda, no cron job, no extra cost

Split sensitive data to protect it from TTL deletion:

◈ DIAGRAM
Sessions table (TTL enabled):
UserId SessionToken ExpTime
user-001 token-abc123 1704067200 ← auto-deleted after expiry
Users table (no TTL):
UserId Name Email City
user-001 Rahul rahul@devops.in Mumbai ← never deleted

Session expires → only the Sessions row deletes. User profile stays safe in the Users table.

Backups for Disaster Recovery

PITR — Point In Time Recovery (continuous backup):

TEXT
AWS continuously records every change behind the scenes
Restore to any specific second within the last 35 days
Must be explicitly enabled — not on by default
Recovery creates a new table — original untouched
Good for: accidental bulk deletes, data corruption, time-travel debugging

On-Demand Backups:

TEXT
Full backup at any time on demand
Kept until you explicitly delete — no automatic expiry
No impact on table performance during backup
Configurable through AWS Backup for cross-region copy
Good for: long-term retention, compliance, pre-release snapshots

S3 Export — analytics without consuming RCU:

TEXT
Export table data to S3 at any point within the PITR retention window
Does not consume any read capacity from the live table
Exported in DynamoDB JSON or ION format
Query with Athena without touching production table
Bash
## Export DynamoDB table to S3 for Athena analytics
aws dynamodb export-table-to-point-in-time \
--table-arn arn:aws:dynamodb:ap-south-1:123456789012:table/Orders \
--s3-bucket devops-analytics-exports \
--s3-prefix dynamodb-exports/ \
--export-format DYNAMODB_JSON \
--region ap-south-1

Hands-on Lab — Create Table, Insert Items, Enable TTL and Streams

Step 1 — Create a DynamoDB table

◈ DIAGRAM
DynamoDB → Tables → Create table
Table name: GameScores
Partition key: UserId Type: String
Sort key: GameId Type: String
Table settings: Default settings
Create table — wait ~30 seconds until Active

Step 2 — Add items

◈ DIAGRAM
DynamoDB → Tables → GameScores → Explore table items → Create item
First item:
UserId: rahul-001 GameId: game-42
Add new attribute → Number → Score: 92
Add new attribute → String → Result: Win
Create item
Second item:
UserId: rahul-001 GameId: game-18 Score: 45 Result: Lose
Both items share UserId = rahul-001 but different GameId.
Partition key + Sort key combination must be unique.
Same user, different games — both valid.

Step 3 — Query for one user

◈ DIAGRAM
GameScores → Explore table items
Switch toggle from Scan → Query
UserId (partition key): rahul-001 → Run
Returns only rahul-001 items efficiently.
Scan reads every item in the table — never use it in production code.

Step 4 — Enable TTL

◈ DIAGRAM
DynamoDB → Tables → GameScores → Additional settings tab
Time to live (TTL) → Enable → Attribute name: ExpTime → Enable TTL
Now create a temporary item:
UserId: rahul-001 GameId: session-temp
Add attribute → Number → ExpTime: (paste a unix timestamp 5 min from now)
Bash
## Get a timestamp 5 minutes from now
date -d "+5 minutes" +%s
TEXT
After 5 minutes DynamoDB deletes this item automatically.
Zero cost. Zero code. Perfect for sessions and temporary records.

Step 5 — Enable DynamoDB Streams

◈ DIAGRAM
DynamoDB → Tables → GameScores → Exports and streams tab
DynamoDB stream details → Enable
View type: New and old images → Enable stream
Every create, update, delete now appears in the stream.
Lambda or Kinesis can read from the stream in real time.

Step 6 — Enable PITR

◈ DIAGRAM
DynamoDB → Tables → GameScores → Backups tab
Point-in-time recovery → Enable
You can now restore to any second in the last 35 days.
Try it: Backups → Restore to point in time → pick any recent timestamp.

Step 7 — Cleanup

◈ DIAGRAM
DynamoDB → Tables → GameScores → Delete
Confirm by typing the table name → Delete table

Production Best Practices and Common Pitfalls

  • Choose a high-cardinality partition key — the partition key determines data distribution and is the single most important table design decision
  • Use Query not Scan — Scan reads every item in the table and filters after, extremely expensive at scale. Design access patterns around Query
  • Enable PITR on every production table from day one — accidental deletes happen and you cannot recover without it
  • Enable Streams before adding Global Tables — it is a hard prerequisite and cannot be added retroactively to an existing Global Table
  • Start with On-Demand capacity for new tables — switch to Provisioned with Auto Scaling once traffic is predictable
  • Store large objects in S3 and save only the S3 URL in DynamoDB — maximum item size is 400 KB, images and documents do not belong in DynamoDB
  • Use TTL for any data with a natural expiry — session tokens, temporary codes, cache entries — zero cost, zero cleanup code needed

Quick Reference and Troubleshooting Commands

Task Command
Create table aws dynamodb create-table --table-name <name> --attribute-definitions ... --key-schema ... --billing-mode PAY_PER_REQUEST
Put item aws dynamodb put-item --table-name <name> --item '{...}'
Get item aws dynamodb get-item --table-name <name> --key '{...}'
Query items aws dynamodb query --table-name <name> --key-condition-expression "PK = :val" --expression-attribute-values '{...}'
Scan table aws dynamodb scan --table-name <name>
Delete item aws dynamodb delete-item --table-name <name> --key '{...}'
Enable TTL aws dynamodb update-time-to-live --table-name <name> --time-to-live-specification "Enabled=true,AttributeName=ExpTime"
Enable PITR aws dynamodb update-continuous-backups --table-name <name> --point-in-time-recovery-specification PointInTimeRecoveryEnabled=true
Enable Streams aws dynamodb update-table --table-name <name> --stream-specification StreamEnabled=true,StreamViewType=NEW_AND_OLD_IMAGES
List tables aws dynamodb list-tables --region ap-south-1

Common problems and fixes:

Problem Likely cause Fix
ProvisionedThroughputExceededException Under-provisioned or hot partition Switch to On-Demand or redesign partition key
Item too large Item exceeds 400 KB Store large attributes in S3, save URL in DynamoDB
Global Table creation fails Streams not enabled Enable Streams first with NEW_AND_OLD_IMAGES
Query returns no results Wrong key condition or expression attribute names Verify partition key value matches exactly, check expression syntax
Scan getting throttled Scan consuming all read capacity Replace with Query using proper key design, or add GSI
TTL items not deleting TTL attribute not set to epoch seconds Verify attribute is a Number type containing Unix epoch timestamp
Common Mistake

Using Scan instead of Query. Scan reads every single item in the table and then filters results in memory. On a 10 million item table this consumes enormous read capacity, takes seconds, and gets progressively worse as the table grows. Always design your access patterns around Query using the primary key. If you find yourself writing Scans in production code, your table design needs rethinking.

Common Mistake

Choosing DynamoDB for everything. DynamoDB is excellent at scale with simple access patterns. It is terrible for ad-hoc reporting, complex joins, and ACID transactions across multiple entities. A Zerodha-style trading ledger needs relational integrity — DynamoDB cannot replace that. Always choose the database based on the access pattern, not familiarity.

Security

DynamoDB Global Tables use last-writer-wins for conflict resolution. If two users in different regions update the same item within milliseconds of each other, the last write wins and the other is lost. Design your application to avoid concurrent writes to the same item from multiple regions, or use conditional writes to detect and handle conflicts explicitly.

Common Mistakes to Avoid

Common Mistake

Using Scan instead of Query. Scan reads every single item in the table and then filters. On a 10-million-item table this is slow and expensive. Design access patterns around Query using the primary key from day one.

Common Mistake

Choosing a low-cardinality partition key like Status or Country. If 90% of writes have Status=active, 90% of writes hit one partition. DynamoDB throttles it. Use a high-cardinality key like UserId or OrderId.

Common Mistake

Choosing DynamoDB for everything. DynamoDB is excellent for simple key-based access at massive scale. It is wrong for complex joins, relational integrity, and ad-hoc reporting. Choose the database based on the access pattern.

Resources

AWS Direct Connect vs Site-to-Site VPN Failover

AWS Direct Connect vs Site-to-Site VPN Failover

Direct Connect vs VPN isn't really either/or for production — it's a primary-plus-failover pattern. Here's how to design it, and when either/or is right.

5 min read•Aug 2026
Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot, at the crossover where Lambda stops being cheaper — 2026 pricing, invocation thresholds, and interruption math.

5 min read•Aug 2026
Secrets Manager vs Parameter Store vs Vault

Secrets Manager vs Parameter Store vs Vault

AWS Secrets Manager, Parameter Store, and HashiCorp Vault compared for 2026 - cost math, rotation, multi-cloud fit, and the Vault-to-OpenBao fork.

5 min read•Aug 2026
AWS VPC Security: Hardening Every Layer

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

5 min read•Jul 2026
Event-Driven Architecture on AWS Explained

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

5 min read•Jul 2026
S3 vs RDS vs DynamoDB: Choosing AWS Storage

S3 vs RDS vs DynamoDB: Choosing AWS Storage

Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.

5 min read•Jul 2026
AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

5 min read•Jul 2026
EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2, Lambda, or Fargate — choosing the wrong AWS compute option costs you money and performance. Here is exactly when to use each one in production.

5 min read•Jul 2026

Explore More in AWS Storage and Databases

All 6 Topics

Frequently Asked Questions

Is Amazon DynamoDB - Serverless NoSQL at Scale free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Amazon DynamoDB - Serverless NoSQL at Scale topic cover?

Design DynamoDB tables with the right primary key, capacity mode, and access patterns, and use Streams, Global Tables, DAX, and TTL for production workloads.