What you will learn
- What DynamoDB is and when NoSQL is the right choice over relational databases
- Tables, items, attributes — and how flexible schema actually works in practice
- Primary key design — Partition Key only vs Partition Key + Sort Key
- Why partition key selection determines the performance ceiling of your entire table
- On-Demand vs Provisioned capacity modes and when each saves money
- DAX — microsecond reads with zero application code changes
- DynamoDB Streams — the ordered changelog that powers event-driven architecture
- Global Tables — active-active multi-region replication and the Streams prerequisite
- TTL — automatic item expiry with zero extra cost
- PITR and on-demand backups for disaster recovery
- S3 export for analytics without consuming read capacity
Why this matters
Hotstar serves metadata for millions of videos — episode details, thumbnails, descriptions. The access pattern is simple: get video by ID. The scale is massive: millions of reads per second with single-digit millisecond response time required. RDS cannot handle this without enormous cost and operational complexity. CRED stores user reward points and transaction history — flexible schema evolves as the product adds new reward types without schema migrations. PhonePe processes UPI transactions where session tokens must expire automatically after 10 minutes — TTL handles this at zero cost. DynamoDB is not a replacement for relational databases. It is the correct choice when the access pattern is simple, the scale is massive, and the schema evolves rapidly.
What is DynamoDB and When to Use It
DynamoDB is AWS's fully managed proprietary NoSQL database. Millisecond latency at any scale, serverless, Multi-AZ by default. No schema enforcement — each item can have different attributes.
When DynamoDB wins: Data does not need joins across multiple tables Schema evolves frequently — agile development, new features weekly Need single-digit millisecond reads at massive scale Serverless — no capacity planning needed (On-Demand mode) Need global multi-region active-active writes (Global Tables) Session data, tokens, temporary records with automatic expiry (TTL) When DynamoDB loses: You need SQL joins across multiple entities Complex reporting queries with GROUP BY, aggregations ACID transactions across many tables (use RDS instead) Ad-hoc queries on arbitrary fields you have not indexedDynamoDB vs RDS — when to choose which:
| Question | If Yes → Use |
|---|---|
| Do your entities have complex relationships? | RDS |
| Do you need SQL joins in production queries? | RDS |
| Do you need complex aggregations and reporting? | RDS |
| Is your scale massive with simple key-based access? | DynamoDB |
| Does your schema change frequently? | DynamoDB |
| Do you need global multi-region active-active writes? | DynamoDB |
| Do you need millisecond latency at any scale serverlessly? | DynamoDB |
Tables, Items, Attributes, and Primary Keys
Key terms:
Table → named container for related data, like a database table No schema enforcement — only rule is every item must have the Primary KeyItem → one record, like a row in SQL Maximum size per item: 400 KBAttribute → a field of an item, like a column Different items can have completely different attributesFlexible attributes in practice:
RDS — every row must have the same columns:id name email age1 Rahul rahul@devops.in 252 Priya priya@devops.in NULL ← age must exist even if null DynamoDB — each item has only what it needs:Item 1: { id: "1", name: "Rahul", email: "rahul@devops.in", age: 25 }Item 2: { id: "2", name: "Priya", phone: "9876543210", city: "Pune" }Item 3: { id: "3", name: "Arjun", score: 99, level: 5, badges: ["gold","silver"] }Item 2 has phone and city but no email. Item 3 has score, level, and a list of badges. All in the same table, no errors. This is why DynamoDB enables rapid schema evolution — new attributes just appear on new items.
Data types:
Scalar: String, Number, Binary, Boolean, NullDocument: List, Map (nested structures like JSON objects)Set: String Set, Number Set, Binary SetPrimary Key — decided at table creation, cannot be changed:
Two options:
Option 1 — Partition Key only (simple primary key): One attribute must be unique per item Example: UserId — each user has one item Option 2 — Partition Key + Sort Key (composite primary key): Combination of two attributes must be unique Example: UserId + GameId — same user can have many game scoresGame scores table example:
Partition Key Sort Key Score Result(UserId) (GameId)user-001 game-42 92 Winuser-002 game-18 14 Loseuser-002 game-42 77 Win ← same user, different gameUser-002 appears twice — same user played two games. The UserId + GameId combination is unique so both rows are valid. Sort Key allows one entity to have many related items in the same table.
Partition Key Design — The Most Critical Decision
The partition key determines which physical partition holds your data. A poor choice causes hot partitions — one partition gets all the traffic while others sit idle. This is the most common DynamoDB performance mistake.
Good partition keys (high cardinality, even distribution): UserId → millions of unique users, traffic spread evenly OrderId → unique per order, completely random distribution DeviceId → millions of IoT devices, random access pattern SessionToken → random UUID, perfectly distributed Bad partition keys (low cardinality, hot partitions): Status → only a few values (active, inactive) — all active users on one partition Country → only a few hundred values — India has far more traffic than others Boolean flag → only two values — true gets all writes if most items are true Date → today's date is all writes — yesterday's date gets nothingCommon MistakeUsing a low-cardinality value like Status or Country as the partition key. If 90% of your writes have Status=active, then 90% of your writes go to a single partition. DynamoDB throttles that partition. Your application sees errors. Fix by choosing a high-cardinality attribute like UserId or using partition key sharding.
Read and Write Capacity Modes
On-Demand Mode:
DynamoDB automatically scales reads and writes up and down instantly. No capacity planning needed.
Pay only for what you actually useHandles any traffic without throttling or configurationMore expensive per request than ProvisionedBest for: unpredictable traffic, new apps with unknown usage, spiky workloadsProvisioned Mode:
You specify upfront how many reads per second (RCU) and writes per second (WCU).
RCU = Read Capacity Unit = one strongly consistent read per second for items up to 4 KBWCU = Write Capacity Unit = one write per second for items up to 1 KBCan add Auto Scaling to adjust RCU and WCU automatically within set limitsCheaper per request than On-Demand for predictable steady workloadsRisk: under-provision → throttling; over-provision → wasted cost| Provisioned Mode | On-Demand Mode | |
|---|---|---|
| Capacity planning | Required | None |
| Cost per request | Cheaper | More expensive |
| Throttling risk | Yes if under-provisioned | No |
| Auto Scaling | Optional add-on | Built-in automatic |
| Best for | Predictable steady traffic | Unpredictable or spiky traffic |
TipStart with On-Demand for new applications — no guessing required. Once you understand your traffic pattern after a few weeks, switch to Provisioned with Auto Scaling to save cost.
DAX — DynamoDB Accelerator
DynamoDB responds in single-digit milliseconds. For read-heavy applications hitting the same data constantly — a leaderboard, product catalog, popular items — those milliseconds add up and the database gets congested with repeated reads.
DAX is a fully managed in-memory cache specifically for DynamoDB. Frequently accessed data is returned in microseconds — up to 10x faster than going to DynamoDB directly.
Without DAX: Every request → DynamoDB → single-digit milliseconds 1000 requests for same product → 1000 DynamoDB reads With DAX: First request → DAX miss → DynamoDB → cached in DAX Next 999 requests → DAX hit → microseconds, DynamoDB never touchedNo application code changes needed:
DAX uses the same DynamoDB API. You point your SDK at the DAX cluster endpoint instead of the DynamoDB endpoint. Everything else stays the same.
Old endpoint: dynamodb.ap-south-1.amazonaws.comNew endpoint: my-dax-cluster.abc123.dax-clusters.ap-south-1.amazonaws.comDAX vs ElastiCache for DynamoDB:
DAX: Purpose-built for DynamoDB Zero code changes — same API Caches individual items and query results automatically Use when: repeated reads of the same DynamoDB items ElastiCache: General purpose cache Requires code changes — check cache first manually Caches computed results, aggregations Use when: need to cache complex computations or results DynamoDB cannot produceTipFor DynamoDB read caching, always try DAX first. It requires no code changes and gives microsecond reads. Only reach for ElastiCache if you need to cache something DAX cannot — like an aggregated result computed across many items.
DynamoDB Streams
Every time an item is created, updated, or deleted, DynamoDB Streams captures that change as an ordered event. Think of it as a real-time changelog of everything that happened in your table.
Application creates new user in DynamoDB ↓DynamoDB Streams captures the INSERT event ↓Lambda triggered by the stream ↓Lambda sends welcome email via SES ↓No polling, no scheduled jobs — fully event-drivenTwo ways to consume streams:
DynamoDB Streams (original): 24-hour retention Limited consumers Process with Lambda or Kinesis Adapter Good for simple real-time triggers Kinesis Data Streams (newer, more powerful): 1-year retention High number of consumers Process with Lambda, Kinesis Analytics, Firehose, Glue Streaming Good for complex pipelines, analytics, long-term archivingFull stream processing pattern:
Application writes to DynamoDB ↓DynamoDB Streams ├── Lambda → send notification, update derived table ├── Kinesis Data Streams → Firehose → S3 (archive) └── Kinesis Data Streams → Redshift (analytics)Global Tables — Active-Active Multi-Region
Global Tables makes your DynamoDB table available in multiple regions simultaneously with active-active replication — writes accepted in any region, replicated everywhere.
ap-south-1 (Mumbai) us-east-1 (Virginia)User in India writes User in US writesto Mumbai replica ←→ to Virginia replica ↓ ↓Change replicates Change replicatesto Virginia to Mumbaiwithin seconds within secondsBoth regions are fully active. A user in Mumbai reads from and writes to the Mumbai replica with single-digit millisecond latency. A user in New York does the same against the Virginia replica.
DynamoDB Streams must be enabled before creating Global Tables:
Global Tables uses Streams internally as the replication engine — every change recorded in Streams is what gets replicated to other regions.
Create table → Enable DynamoDB Streams → Create Global Table → Add regionsRememberDynamoDB Streams is a hard prerequisite for Global Tables. If you try to create a Global Table without Streams enabled, it fails. Enable Streams first, always.
Global Tables vs RDS Cross-Region Read Replica:
| DynamoDB Global Tables | RDS Cross-Region Read Replica | |
|---|---|---|
| Write regions | All regions (active-active) | One primary only (active-passive) |
| Replication lag | Sub-second | Minutes |
| Conflicts | Last-writer-wins | Not applicable |
| Use case | Global apps needing local writes | Read scaling in another region |
TTL — Time To Live
TTL automatically deletes items from your table after a set time. No cleanup code, no scheduled jobs, no extra cost.
How it works:
You choose any attribute name — ExpTime, TTL, ExpiresAt, anythingTell DynamoDB once in table settings: this is my TTL attributeWhen creating an item, set that attribute to an epoch timestampDynamoDB scans the attribute vs current timeWhen current time passes the timestamp → item automatically deletedFree — does not consume write capacitySession management example:
User logs in → create session item with ExpTime = now + 86400 (24 hours)User active → session item exists, valid24 hours pass → DynamoDB automatically deletes the session itemNo cleanup Lambda, no cron job, no extra costSplit sensitive data to protect it from TTL deletion:
Sessions table (TTL enabled):UserId SessionToken ExpTimeuser-001 token-abc123 1704067200 ← auto-deleted after expiry Users table (no TTL):UserId Name Email Cityuser-001 Rahul rahul@devops.in Mumbai ← never deletedSession expires → only the Sessions row deletes. User profile stays safe in the Users table.
Backups for Disaster Recovery
PITR — Point In Time Recovery (continuous backup):
AWS continuously records every change behind the scenesRestore to any specific second within the last 35 daysMust be explicitly enabled — not on by defaultRecovery creates a new table — original untouchedGood for: accidental bulk deletes, data corruption, time-travel debuggingOn-Demand Backups:
Full backup at any time on demandKept until you explicitly delete — no automatic expiryNo impact on table performance during backupConfigurable through AWS Backup for cross-region copyGood for: long-term retention, compliance, pre-release snapshotsS3 Export — analytics without consuming RCU:
Export table data to S3 at any point within the PITR retention windowDoes not consume any read capacity from the live tableExported in DynamoDB JSON or ION formatQuery with Athena without touching production table## Export DynamoDB table to S3 for Athena analyticsaws dynamodb export-table-to-point-in-time \ --table-arn arn:aws:dynamodb:ap-south-1:123456789012:table/Orders \ --s3-bucket devops-analytics-exports \ --s3-prefix dynamodb-exports/ \ --export-format DYNAMODB_JSON \ --region ap-south-1Hands-on Lab — Create Table, Insert Items, Enable TTL and Streams
Step 1 — Create a DynamoDB table
DynamoDB → Tables → Create tableTable name: GameScoresPartition key: UserId Type: StringSort key: GameId Type: StringTable settings: Default settingsCreate table — wait ~30 seconds until ActiveStep 2 — Add items
DynamoDB → Tables → GameScores → Explore table items → Create item First item:UserId: rahul-001 GameId: game-42Add new attribute → Number → Score: 92Add new attribute → String → Result: WinCreate item Second item:UserId: rahul-001 GameId: game-18 Score: 45 Result: Lose Both items share UserId = rahul-001 but different GameId.Partition key + Sort key combination must be unique.Same user, different games — both valid.Step 3 — Query for one user
GameScores → Explore table itemsSwitch toggle from Scan → QueryUserId (partition key): rahul-001 → Run Returns only rahul-001 items efficiently.Scan reads every item in the table — never use it in production code.Step 4 — Enable TTL
DynamoDB → Tables → GameScores → Additional settings tabTime to live (TTL) → Enable → Attribute name: ExpTime → Enable TTL Now create a temporary item:UserId: rahul-001 GameId: session-tempAdd attribute → Number → ExpTime: (paste a unix timestamp 5 min from now)## Get a timestamp 5 minutes from nowdate -d "+5 minutes" +%sAfter 5 minutes DynamoDB deletes this item automatically.Zero cost. Zero code. Perfect for sessions and temporary records.Step 5 — Enable DynamoDB Streams
DynamoDB → Tables → GameScores → Exports and streams tabDynamoDB stream details → EnableView type: New and old images → Enable stream Every create, update, delete now appears in the stream.Lambda or Kinesis can read from the stream in real time.Step 6 — Enable PITR
DynamoDB → Tables → GameScores → Backups tabPoint-in-time recovery → Enable You can now restore to any second in the last 35 days.Try it: Backups → Restore to point in time → pick any recent timestamp.Step 7 — Cleanup
DynamoDB → Tables → GameScores → DeleteConfirm by typing the table name → Delete tableProduction Best Practices and Common Pitfalls
- Choose a high-cardinality partition key — the partition key determines data distribution and is the single most important table design decision
- Use Query not Scan — Scan reads every item in the table and filters after, extremely expensive at scale. Design access patterns around Query
- Enable PITR on every production table from day one — accidental deletes happen and you cannot recover without it
- Enable Streams before adding Global Tables — it is a hard prerequisite and cannot be added retroactively to an existing Global Table
- Start with On-Demand capacity for new tables — switch to Provisioned with Auto Scaling once traffic is predictable
- Store large objects in S3 and save only the S3 URL in DynamoDB — maximum item size is 400 KB, images and documents do not belong in DynamoDB
- Use TTL for any data with a natural expiry — session tokens, temporary codes, cache entries — zero cost, zero cleanup code needed
Quick Reference and Troubleshooting Commands
| Task | Command |
|---|---|
| Create table | aws dynamodb create-table --table-name <name> --attribute-definitions ... --key-schema ... --billing-mode PAY_PER_REQUEST |
| Put item | aws dynamodb put-item --table-name <name> --item '{...}' |
| Get item | aws dynamodb get-item --table-name <name> --key '{...}' |
| Query items | aws dynamodb query --table-name <name> --key-condition-expression "PK = :val" --expression-attribute-values '{...}' |
| Scan table | aws dynamodb scan --table-name <name> |
| Delete item | aws dynamodb delete-item --table-name <name> --key '{...}' |
| Enable TTL | aws dynamodb update-time-to-live --table-name <name> --time-to-live-specification "Enabled=true,AttributeName=ExpTime" |
| Enable PITR | aws dynamodb update-continuous-backups --table-name <name> --point-in-time-recovery-specification PointInTimeRecoveryEnabled=true |
| Enable Streams | aws dynamodb update-table --table-name <name> --stream-specification StreamEnabled=true,StreamViewType=NEW_AND_OLD_IMAGES |
| List tables | aws dynamodb list-tables --region ap-south-1 |
Common problems and fixes:
| Problem | Likely cause | Fix |
|---|---|---|
| ProvisionedThroughputExceededException | Under-provisioned or hot partition | Switch to On-Demand or redesign partition key |
| Item too large | Item exceeds 400 KB | Store large attributes in S3, save URL in DynamoDB |
| Global Table creation fails | Streams not enabled | Enable Streams first with NEW_AND_OLD_IMAGES |
| Query returns no results | Wrong key condition or expression attribute names | Verify partition key value matches exactly, check expression syntax |
| Scan getting throttled | Scan consuming all read capacity | Replace with Query using proper key design, or add GSI |
| TTL items not deleting | TTL attribute not set to epoch seconds | Verify attribute is a Number type containing Unix epoch timestamp |
Common MistakeUsing Scan instead of Query. Scan reads every single item in the table and then filters results in memory. On a 10 million item table this consumes enormous read capacity, takes seconds, and gets progressively worse as the table grows. Always design your access patterns around Query using the primary key. If you find yourself writing Scans in production code, your table design needs rethinking.
Common MistakeChoosing DynamoDB for everything. DynamoDB is excellent at scale with simple access patterns. It is terrible for ad-hoc reporting, complex joins, and ACID transactions across multiple entities. A Zerodha-style trading ledger needs relational integrity — DynamoDB cannot replace that. Always choose the database based on the access pattern, not familiarity.
SecurityDynamoDB Global Tables use last-writer-wins for conflict resolution. If two users in different regions update the same item within milliseconds of each other, the last write wins and the other is lost. Design your application to avoid concurrent writes to the same item from multiple regions, or use conditional writes to detect and handle conflicts explicitly.
Common Mistakes to Avoid
Common MistakeUsing Scan instead of Query. Scan reads every single item in the table and then filters. On a 10-million-item table this is slow and expensive. Design access patterns around Query using the primary key from day one.
Common MistakeChoosing a low-cardinality partition key like Status or Country. If 90% of writes have Status=active, 90% of writes hit one partition. DynamoDB throttles it. Use a high-cardinality key like UserId or OrderId.
Common MistakeChoosing DynamoDB for everything. DynamoDB is excellent for simple key-based access at massive scale. It is wrong for complex joins, relational integrity, and ad-hoc reporting. Choose the database based on the access pattern.