What you will learn
- Why caching exists and the two database problems it solves
- Redis vs Memcached — the feature comparison that drives every selection decision
- Lazy Loading — cache-aside pattern, cache misses, and the staleness trade-off
- Write-Through — update cache on every write, zero staleness but higher cost
- Session Store — stateless EC2 behind a load balancer without sticky sessions
- ElastiCache for Redis vs DAX — which one for DynamoDB, which for RDS
- ElastiCache Serverless — auto-scaling cache without cluster sizing
- Redis Cluster mode — sharding data across multiple primaries for scale
- Multi-AZ with automatic failover for production high availability
- How to connect to ElastiCache from within the same VPC
Why this matters
At peak on Swiggy, the same restaurant menu is requested by thousands of users simultaneously. Without caching, each request hits the RDS database — thousands of identical SELECT queries per second. The database becomes the bottleneck. Response times climb. The database CPU maxes out. With ElastiCache Redis sitting in front of RDS, the first request fetches from the database and caches the result. The next ten thousand requests hit Redis — served in under a millisecond, database never touched. This is the core value of caching. Understanding when to use Lazy Loading versus Write-Through, and why Session Store eliminates the need for sticky sessions, is what lets you build applications that scale beyond what a single database can serve.
Why Caching Exists
Every production application has two database problems at scale:
Problem 1 — Repeated reads: Same data requested thousands of times per minute Same query hitting the database identically Database CPU climbs, latency increases for ALL queries Problem 2 — Database read latency: Even the fastest RDS query takes 5-20ms At scale, 5ms per request × 100,000 requests/sec = significant overhead Both problems solved by a cache: First request: check cache → miss → query DB → store in cache → return Next N requests: check cache → HIT → return in <1ms → DB never touchedElastiCache is AWS's managed in-memory caching service. Data lives in RAM — microsecond latency, no disk I/O, no SQL parsing overhead.
Redis vs Memcached
Both Redis and Memcached are in-memory caching engines. The choice comes down to features.
| Feature | Redis | Memcached |
|---|---|---|
| Data structures | Strings, Hashes, Lists, Sets, Sorted Sets, Bitmaps, HyperLogLog | Strings only |
| Persistence | Yes — RDB snapshots, AOF log | No — memory only |
| Replication | Yes — primary with read replicas | No |
| Multi-AZ failover | Yes — automatic | No |
| Pub/Sub messaging | Yes | No |
| Transactions | Yes (MULTI/EXEC) | No |
| Lua scripting | Yes | No |
| Cluster sharding | Yes — Redis Cluster mode | Yes — multi-threaded simple sharding |
| Use case | Most production workloads | Simple multi-threaded cache with no persistence |
When to choose Memcached:
You need multi-threading for extreme throughputYou need simple horizontal scaling by adding more nodesYou do not need persistence, replication, or pub/subData structures beyond strings are never neededWhen to choose Redis (almost always):
You need High Availability with automatic failoverYou need data persistence (cache survives restart)You use complex data structures (leaderboards with Sorted Sets)You need pub/sub messagingYou need geospatial queriesYou need session storage with expiryRememberRedis supports replication and multi-AZ failover. Memcached does not. For any production workload where the cache going down would cause a problem, use Redis — its primary can fail and a replica promotes automatically.
Caching Pattern 1 — Lazy Loading (Cache-Aside)
The most common caching pattern. The application manages the cache explicitly.
Read path: Application checks cache: GET user:123 Cache HIT → return data immediately (fast, <1ms) Cache MISS → query database → store result in cache → return to user Write path: Update database Invalidate or do NOT update the cache (next read will be a miss → fetches fresh data from DB → repopulates cache)import redisimport json cache = redis.Redis(host='devops-cache.abc123.ng.0001.aps1.cache.amazonaws.com', port=6379)db = get_db_connection() def get_restaurant_menu(restaurant_id): cache_key = f"menu:{restaurant_id}" ## Step 1: Check cache first cached = cache.get(cache_key) if cached: return json.loads(cached) ## Cache HIT — return immediately ## Step 2: Cache miss — query database menu = db.query("SELECT * FROM menus WHERE restaurant_id = %s", restaurant_id) ## Step 3: Store in cache with TTL of 300 seconds cache.setex(cache_key, 300, json.dumps(menu)) return menuLazy Loading advantages:
Only requested data is cached — no wasted memory for unused dataCache failure does not break the application — falls back to DBSimple to implementLazy Loading disadvantages:
First request after cache miss is slow — hits the DBCache can contain stale data until TTL expiresCache stampede: if cache is empty and 1000 users request the same key simultaneously,all 1000 hit the database before the first one repopulates the cacheCache stampede prevention:
## Use a lock to prevent stampede — only one request populates cacheimport redis_lock def get_restaurant_menu(restaurant_id): cache_key = f"menu:{restaurant_id}" cached = cache.get(cache_key) if cached: return json.loads(cached) ## Acquire lock — other requests wait instead of hitting DB with redis_lock.Lock(cache, f"lock:{cache_key}", expire=10): ## Check again after acquiring lock — another thread may have populated cached = cache.get(cache_key) if cached: return json.loads(cached) menu = db.query("SELECT * FROM menus WHERE restaurant_id = %s", restaurant_id) cache.setex(cache_key, 300, json.dumps(menu)) return menuCaching Pattern 2 — Write-Through
Update the cache every time you write to the database. Cache and database always in sync.
Write path: Application writes to database Application immediately updates cache with same data Cache always has current data Read path: Application checks cache Almost always a HIT — cache has everything because all writes update it Very few cache missesdef update_restaurant_menu(restaurant_id, menu_data): ## Step 1: Write to database db.execute("UPDATE menus SET data = %s WHERE restaurant_id = %s", json.dumps(menu_data), restaurant_id) ## Step 2: Update cache immediately — same TTL as before cache_key = f"menu:{restaurant_id}" cache.setex(cache_key, 3600, json.dumps(menu_data)) return TrueWrite-Through advantages:
Cache always has current data — no stalenessReads are almost always cache HITs — fast for everyoneWrite-Through disadvantages:
Every write hits both database AND cache — two operations per writeCache filled with data that may never be read (wasteful memory use)Write penalty — slightly slower writes because both must succeedRememberLazy Loading = fast writes, risk of stale reads. Write-Through = fast reads, slower writes. In practice, combine them: Write-Through for frequently read frequently written data, Lazy Loading for read-heavy data that changes rarely.
Caching Pattern 3 — Session Store
The problem: a user logs in and lands on EC2 Instance A. Their session (cart, auth token, preferences) is stored in that instance's memory. The load balancer sends their next request to EC2 Instance B. Session is gone. User sees a login page.
Wrong fix — sticky sessions:
Force the load balancer to always route the same user to the same EC2 instance. Causes uneven load distribution. If that instance fails, the session is lost anyway.
Right fix — Session Store with Redis:
Store all sessions in Redis, not in EC2 memory. Any instance can serve any user because sessions live outside the instances.
User logs in: Application creates session → stores in Redis with 24h TTL Returns session token to user's browser (cookie) Next request (any EC2 instance can handle): User sends request with session token EC2 reads session from Redis → gets user data Serves the request No sticky sessions. Any instance serves any user. Fully scalable.from flask import Flask, sessionfrom flask_session import Sessionimport redis app = Flask(__name__) ## Store sessions in ElastiCache Redis — not in EC2 memoryapp.config['SESSION_TYPE'] = 'redis'app.config['SESSION_REDIS'] = redis.Redis( host='devops-cache.abc123.ng.0001.aps1.cache.amazonaws.com', port=6379)app.config['PERMANENT_SESSION_LIFETIME'] = 86400 ## 24 hours TTL Session(app) def login(): session['user_id'] = authenticate_user(request.form) session['cart'] = [] return redirect('/dashboard') def dashboard(): if 'user_id' not in session: return redirect('/login') ## Session retrieved from Redis automatically — works from any EC2 instance return render_template('dashboard.html', user_id=session['user_id'])ElastiCache Redis vs DAX — Which for DynamoDB
Both are in-memory caches. The question is which to use when your backend is DynamoDB.
DAX (DynamoDB Accelerator): Purpose-built for DynamoDB Zero code changes — same API, just change the endpoint Caches individual GetItem and Query results automatically Microsecond latency for read-heavy DynamoDB workloads Cannot cache computed aggregations or complex results ElastiCache Redis with DynamoDB: General purpose cache Requires code changes — check cache first, fall back to DynamoDB Can cache anything — aggregations, processed results, custom keys Pub/sub, sorted sets, complex data structures available Decision: Simple DynamoDB read caching → DAX (zero code changes) Need to cache complex computations or non-DynamoDB data → ElastiCache RedisElastiCache Architecture — Cluster Modes
Cluster Mode Disabled (single primary + replicas):
Primary node: handles all reads and writesRead Replicas: serve read traffic only (up to 5 replicas)Multi-AZ: one replica in different AZ promoted on primary failure Primary (ap-south-1a) → Replica 1 (ap-south-1b) → Replica 2 (ap-south-1c) One shard — all data fits in one primary node's memoryGood for: small to medium datasets, read-heavy workloadsCluster Mode Enabled (multiple shards):
Data partitioned across multiple primary nodes (shards)Each shard has its own set of replicasScales beyond single node memory limits Shard 1 (keys A-M): Primary + 2 ReplicasShard 2 (keys N-Z): Primary + 2 Replicas Up to 500 nodes, 90+ shardsScales both writes AND reads (unlike single-primary which only scales reads)Good for: large datasets, write-heavy workloads, truly large-scale productionAutomatic failover:
Primary fails ↓ElastiCache detects failure (30-60 seconds) ↓Promotes a replica to primary automatically ↓DNS endpoint updated to point to new primary ↓Application reconnects to new primary automatically(application must handle reconnection — connection pooling helps)ElastiCache Serverless
New option — no cluster sizing, no node type selection, no shard planning. You just use it and pay for data stored and compute used.
You provision: nothingAWS handles: scaling up and down automaticallyBilling: per GB stored + per ECU (ElastiCache Compute Unit)Use when:
New applications with unknown cache hit rates and data volumesSpiky or unpredictable workloadsYou want zero capacity planning overheadHands-on Lab — Create ElastiCache Redis and Connect from EC2
Step 1 — Create a Security Group for ElastiCache
EC2 → Security Groups → Create security groupName: redis-sg VPC: your VPCInbound rule: Custom TCP Port: 6379Source: App EC2 Security Group (not 0.0.0.0/0 — Redis must never be public)Create security groupStep 2 — Create an ElastiCache Subnet Group
ElastiCache → Subnet groups → Create subnet groupName: redis-subnet-groupVPC: your VPCSubnets: select your private subnets in 2 AZsCreateStep 3 — Create Redis Cluster
ElastiCache → Redis OSS caches → Create Redis OSS cacheDeployment option: Design your own cacheCreation method: Easy createLocation: AWS CloudName: devops-redisNode type: cache.t3.micro (free tier eligible)Subnet group: redis-subnet-groupSecurity group: redis-sgCreate Wait 3-5 minutes until Status: AvailableStep 4 — Get the endpoint
ElastiCache → Redis OSS caches → devops-redisPrimary endpoint: devops-redis.abc123.cache.amazonaws.com:6379Step 5 — Connect from EC2 and test
SSH into your EC2 instance (must be in the same VPC):## Install redis-clisudo yum install -y redis ## Connect to ElastiCacheredis-cli -h devops-redis.abc123.cache.amazonaws.com -p 6379 ## Test basic operationsSET user:101:name "Rahul Sharma"GET user:101:name## Returns: "Rahul Sharma" SET user:101:city "Mumbai"GET user:101:city## Returns: "Mumbai" ## Set with TTL (expires in 30 seconds — TTL-based session)SET session:abc123 "active" EX 30TTL session:abc123## Returns: 29 (seconds remaining) ## After 30 secondsGET session:abc123## Returns: nil (expired automatically)Step 6 — Test Lazy Loading pattern
## Simulate cache miss → DB lookup → cache store## (In production your app does this, here we simulate manually) ## 1. Check cache — missGET product:P001## Returns: nil (not in cache) ## 2. "Query database" (simulated)## 3. Store in cache with 5 minute TTLSET product:P001 '{"name":"Biryani","price":250}' EX 300 ## 4. Next request — hitGET product:P001## Returns immediately from Redis: {"name":"Biryani","price":250}Step 7 — Cleanup
ElastiCache → Redis OSS caches → devops-redis → Delete (no final backup needed for lab)ElastiCache → Subnet groups → redis-subnet-group → DeleteEC2 → Security Groups → redis-sg → DeleteProduction Best Practices and Common Pitfalls
- Always enable Multi-AZ with automatic failover for production Redis — a single primary with no replica means cache failure takes the entire cache offline
- Use connection pooling in your application — creating a new Redis connection per request is expensive, reuse connections
- Set appropriate TTL on every cached item — cache without TTL grows unbounded and eventually evicts important data
- Use ElastiCache inside the same VPC as your application — never expose Redis to the public internet
- Enable at-rest and in-transit encryption — Redis may hold session tokens and sensitive business data
- Monitor cache hit ratio in CloudWatch — below 80% means the cache is not effective, review your caching strategy
- Use Redis Cluster mode (sharding) when your dataset exceeds single node memory — Cluster mode scales both memory and write throughput
Quick Reference and Troubleshooting Commands
| Task | Command |
|---|---|
| List clusters | aws elasticache describe-replication-groups --region ap-south-1 |
| Get primary endpoint | aws elasticache describe-replication-groups --replication-group-id <id> --query 'ReplicationGroups[0].NodeGroups[0].PrimaryEndpoint' |
| List cache subnet groups | aws elasticache describe-cache-subnet-groups --region ap-south-1 |
| Describe events | aws elasticache describe-events --region ap-south-1 |
| Test Redis connection | redis-cli -h <endpoint> -p 6379 PING |
| Set a key | redis-cli -h <endpoint> SET key value |
| Get a key | redis-cli -h <endpoint> GET key |
| Set key with TTL | redis-cli -h <endpoint> SETEX key 300 value |
| Check remaining TTL | redis-cli -h <endpoint> TTL key |
| List all keys (dev only) | redis-cli -h <endpoint> KEYS '*' |
Common problems and fixes:
| Problem | Likely cause | Fix |
|---|---|---|
| Cannot connect to ElastiCache | Security Group not allowing port 6379 | Add inbound rule allowing 6379 from application SG |
| Cache hit rate below 50% | TTL too short or keys not aligned with query patterns | Increase TTL, review what data is being cached vs what is being queried |
| Memory usage growing unboundedly | No TTL on cached items, eviction not configured | Set TTL on all items, configure maxmemory-policy to allkeys-lru |
| Primary failover not happening | Automatic failover not enabled | Enable automatic-failover-enabled on the replication group |
| Sessions lost between requests | Session data in EC2 memory, not Redis | Implement Session Store pattern with Redis |
Common MistakeRunning a single ElastiCache node with no replicas in production. When that node has an issue, the cache goes offline. Every request falls back to the database simultaneously — the spike of database queries causes the database to also struggle. Always run at least one replica with Multi-AZ failover enabled in production.
TipUse Redis Sorted Sets for leaderboards instead of fetching, sorting, and ranking data from RDS. ZADD adds a player score with O(log N) complexity, ZRANGE retrieves the top N players in order. At Swiggy scale — a leaderboard of 10 million delivery partners ranked by ratings — Redis does this in microseconds. RDS would require a full table sort that takes seconds.
Common Mistakes to Avoid
Common MistakePutting ElastiCache in a public subnet or making it internet-accessible. ElastiCache has no authentication by default (unless enabled). Exposing it to the internet means anyone can read and write your cache. Always place in a private subnet.
Common MistakeThinking ElastiCache works transparently in front of RDS. Your application code must explicitly check the cache first, then fall back to RDS on a miss, then store the result. It is not a proxy layer. It requires code changes.
RememberAlways choose Redis over Memcached for new workloads. Redis supports persistence, replication, Multi-AZ, pub/sub, and rich data structures. Memcached supports none of these. There is almost no production scenario where Memcached wins.