What you will learn
- What serverless means and when Lambda is the right choice over EC2
- Every Lambda trigger and which invocation pattern each uses
- Lambda limits — every number that matters in production
- How Lambda pricing works and why it is almost always the cheapest option for event-driven workloads
- Concurrency, throttling, and how one function can silently starve all others in your account
- Cold starts — what causes them, how bad they actually are, and how to fix them
- Provisioned Concurrency for APIs that cannot tolerate cold start latency
- SnapStart — the free cold start fix for Java functions
- How to put Lambda inside your VPC to reach private RDS and ElastiCache
- RDS Proxy — why Lambda and databases need a connection pool in the middle
- CloudFront Functions vs Lambda@Edge — edge compute at the right level
Why this matters
When a Swiggy user places an order, the notification that goes to the restaurant and the delivery partner does not need a server running 24 hours a day waiting for that event. It needs a function that runs for 200 milliseconds, sends the notification, and stops. Lambda handles that at essentially zero cost.
At Razorpay, Lambda processes webhook events from payment gateways — hundreds of thousands per day, each taking under a second. At CRED, Lambda runs scheduled tasks — computing reward points, generating statements — without maintaining any servers. Serverless is not a buzzword. It is the correct tool for event-driven, short-lived compute at any scale.
Serverless vs EC2 — When Lambda Wins
Serverless means you do not manage, provision, or think about servers. AWS handles everything behind the scenes.
EC2 — always on: Server runs 24/7 whether traffic is zero or massive You pay even at 3 AM when nothing is happening Traffic spike → manually scale or use ASG You manage OS, patching, capacity Lambda — on demand: Event arrives → Lambda runs your function Function finishes → Lambda stops Pay only for the milliseconds it actually ran Scaling is automatic and invisible EC2 is a full-time employee at a desk 24/7, paid whether or not there is work.Lambda is a contractor called in for specific tasks, paid only while working.When Lambda wins:
Short tasks triggered by events (under 15 minutes)Infrequent or unpredictable traffic — pay per request, not per hourFunctions that are idle most of the timeFile processing, notifications, scheduled jobs, API backendsTight cost control at low to medium traffic volumesWhen Lambda loses:
Tasks that run longer than 15 minutesWorkloads needing persistent background processes or websocket connectionsGPU workloads (ML training, video encoding)Very high sustained throughput where Reserved EC2 instances are cheaperLambda Triggers — What Invokes a Function
Lambda runs only when something triggers it. Everything is event-driven.
API Gateway → HTTP request from browser or mobile appALB → HTTP request (alternative to API Gateway)S3 → object uploaded, deleted, or restoredDynamoDB Streams → item created, updated, or deletedKinesis → new record in a data streamSQS → message arrives in a queueSNS → notification published to a topicEventBridge → scheduled event or AWS service eventCloudWatch Logs → log pattern match via subscription filterCognito → user signs up, pre-token generation hookIoT Core → message from a connected deviceTwo invocation types:
Synchronous — caller waits for the result: API Gateway → Lambda → response returned to the caller ALB → Lambda → response returned If Lambda is throttled → 429 ThrottleError returned immediately Caller must handle the error — no automatic retry Asynchronous — caller does not wait: S3 → Lambda → S3 moves on, Lambda processes in background SNS → Lambda → SNS moves on EventBridge → Lambda If throttled → Lambda retries automatically for up to 6 hours Still failing after retries → event sent to Dead Letter QueueLambda Limits — Every Number That Matters
| Limit | Value |
|---|---|
| Maximum execution time | 900 seconds (15 minutes) |
| Memory allocation | 128 MB to 10,240 MB (10 GB) |
| Disk space in /tmp | 512 MB default, up to 10 GB |
| Environment variables | 4 KB total |
| Deployment package compressed | 50 MB |
| Deployment package uncompressed | 250 MB |
| Concurrent executions per account per region | 1,000 (soft limit — request increase) |
| Layers per function | 5 |
RememberIncreasing RAM automatically increases CPU and network bandwidth too. Lambda has no separate CPU control — RAM is the only performance knob. If your function is slow, try doubling the RAM before optimising code. More RAM = more CPU = faster execution = lower duration billing.
Lambda Pricing — Almost Always Cheapest for Event-Driven
Pay per request:
First 1,000,000 requests per month → freeAfter that → $0.20 per million requestsPay per duration (billed in 1ms increments):
First 400,000 GB-seconds per month → freeAfter that → $1.00 per 600,000 GB-seconds GB-seconds = memory in GB × duration in seconds128 MB = 0.125 GBFunction runs for 200ms = 0.2 secondsOne invocation = 0.125 × 0.2 = 0.025 GB-secondsReal example — Swiggy order notification:
Function: 256 MB RAM, runs 200ms per orderTraffic: 1 million orders per day Duration cost:1,000,000 × 0.25 GB × 0.2 sec = 50,000 GB-seconds50,000 / 600,000 × $1.00 = $0.08 per day → about $2.50 per month Same on t3.small EC2 running 24/7 = $14 per monthLambda is 5x cheaper for this workload pattern.Concurrency and Throttling
Concurrency is how many Lambda functions are running at the same moment. If 100 users hit your API simultaneously, Lambda spins up 100 isolated execution environments in parallel.
1 request → 1 Lambda execution environment100 requests → 100 Lambda execution environments simultaneouslyAll finish → all environments idle (may be kept warm for a while)AWS allows up to 1,000 concurrent executions per account per region by default.
The account-level starvation problem:
Without limits, one high-traffic function can consume all 1,000 concurrent slots and starve every other function in the account.
Normal state: Orders API → 200 concurrent Payment API → 300 concurrent Report generator → 100 concurrent Remaining: 400 slots available Traffic spike on Orders API: Orders API → 900 concurrent ← consuming almost everything Payment API → throttled ← payments start failing Report generator → throttled ← reports stop workingReserved Concurrency — fix this:
Set a per-function concurrency limit. Guarantees the function always has capacity AND prevents it from consuming more than its share.
Orders API → reserved: 400 (guaranteed 400 slots, cannot exceed 400)Payment API → reserved: 300 (guaranteed 300 slots, cannot exceed 300)Remaining 300 slots shared by other functions Even during a spike, Orders API cannot take Payment API's slots. Console: Lambda → Function → Configuration → Concurrency → Edit → Reserved concurrencySecurityReserved concurrency also acts as a ceiling — setting it to 0 throttles the function completely. Useful for temporarily disabling a function without deleting it.
Cold Starts — The Hidden Latency
Lambda shuts down execution environments when they are not in use. When a new request arrives and no warm environment is available, Lambda must spin up a new one. This is a cold start.
Cold start — new environment needed: Download your code package Start the runtime (Node.js, Python, Java...) Run initialisation code outside your handler Run your handler function Total: 100ms to 10 seconds depending on runtime and code size Warm start — existing environment reused: Run your handler function only Total: just your function execution timeCold start impact by runtime:
Python, Node.js → 100-500ms cold start (acceptable for most use cases)Java, .NET → 1-10 seconds cold start (serious problem for user-facing APIs)Go, Rust → under 100ms (excellent)Move heavy initialisation outside the handler:
The most impactful code change you can make. Code outside the handler runs once on cold start and is reused across all warm invocations.
## Wrong — database connection created on every invocationdef lambda_handler(event, context): db = connect_to_database() ## runs on EVERY invocation result = db.query("SELECT ...") return result ## Right — connection created once on cold start, reused on warm invocationsdb = connect_to_database() ## runs only on cold start def lambda_handler(event, context): result = db.query("SELECT ...") ## fast every time return resultThis applies to: database connections, SDK clients, configuration loading, secrets fetching — anything that takes time and can be reused.
Provisioned Concurrency — Pre-Warmed Instances
Provisioned Concurrency keeps a set number of execution environments always initialised and ready. Zero cold start for those slots.
Without Provisioned Concurrency: First request → cold start (500ms to 10 seconds) Next requests → warm (just handler time) With Provisioned Concurrency (10 pre-warmed): First 10 simultaneous requests → zero cold start Request 11 and above → may cold start if all 10 are busy Console: Lambda → Function → Configuration → Concurrency → Allocate provisioned concurrency → set number → SaveCost: You pay for provisioned concurrency even when no requests are coming — environments are kept alive and ready.
Use Provisioned Concurrency for:
- User-facing APIs where first-request latency matters
- Payment processing where timeouts cause real failures
- Applications with strict SLA requirements on response time
Combine with Application Auto Scaling:
Scale provisioned concurrency up during business hours and down overnight automatically. You define a schedule or a target utilisation.
SnapStart — Free Cold Start Fix for Java
Provisioned Concurrency solves cold starts but costs extra. SnapStart is free and works for Java (and .NET) functions.
Normal Java cold start (every new environment): Load JVM → load all classes → run init code → run handler Total: 3 to 10 seconds SnapStart — AWS runs init ONCE when you publish a new version: Snapshot of the fully initialised state is saved New cold start restores from snapshot instead of initialising from scratch Total: under 1 second — up to 10x faster, at no extra cost Console: Lambda → Function → Configuration → General configuration → SnapStart → PublishedVersions → Save Then publish a new version to activate SnapStartRememberSnapStart is free. It works per published version. Every time you publish a new version, AWS takes a fresh snapshot after initialisation. There is no reason not to enable it on every Java Lambda function.
Lambda Inside a VPC
By default Lambda runs in an AWS-owned network. Public AWS services like S3, DynamoDB, and public APIs work fine. Private resources — your RDS in a private subnet, your ElastiCache — are unreachable.
Default Lambda: Reaches S3, DynamoDB, public APIs → works Reaches private RDS, ElastiCache → cannotTo reach private resources, deploy Lambda inside your VPC. AWS creates an ENI (Elastic Network Interface) in your private subnet, giving Lambda a private IP inside your VPC.
Lambda with VPC config: Specify: VPC, private subnets, security group Lambda gets a private IP in your subnet Can reach anything in your VPC using private IPs RDS Security Group must allow: Inbound port 3306 from Lambda Security Group IDLambda in VPC and internet access:
Lambda inside a VPC loses access to the public internet by default. If your function needs both — private RDS AND a public external API — add a NAT Gateway.
Lambda (private subnet) → NAT Gateway (public subnet) → internetLambda (private subnet) → private IP → RDS (private subnet)RDS Proxy — Connection Pooling for Lambda
Lambda can scale to 1,000 concurrent executions. Each one opens its own database connection. Under load, that is 1,000 simultaneous connections to your RDS instance. Most databases cannot handle this.
Without RDS Proxy: 1,000 Lambda executions → 1,000 direct DB connections → RDS overwhelmed → errors With RDS Proxy: 1,000 Lambda executions → RDS Proxy → 20 pooled connections → RDS safeRDS Proxy pools and shares connections. Lambda connects to the Proxy. The Proxy manages a small pool of real database connections and multiplexes Lambda requests across them.
Three benefits:
- Scalability — Lambda never overwhelms the database
- Availability — Proxy preserves connections during Multi-AZ failover, reducing failover time by 66%
- Security — enforces IAM authentication, stores credentials in Secrets Manager
RDS Proxy lives inside your VPC. Lambda must also be in the same VPC to reach it.
Common MistakeConnecting Lambda directly to RDS without RDS Proxy when expecting high concurrency. This is fine with 10 concurrent executions. It fails catastrophically at 500. Plan for RDS Proxy from the start for any Lambda-to-database architecture.
CloudFront Functions vs Lambda@Edge
Both run code at CloudFront edge locations — close to the user, not back in your region.
Request from user in Chennai ↓[Viewer Request] ← CloudFront Functions OR Lambda@Edge can run here ↓CloudFront edge cache ↓[Origin Request] ← Lambda@Edge only ↓Your origin (S3, ALB, EC2) ↓[Origin Response] ← Lambda@Edge only ↓CloudFront edge ↓[Viewer Response] ← CloudFront Functions OR Lambda@Edge can run here ↓User in Chennai| CloudFront Functions | Lambda@Edge | |
|---|---|---|
| Language | JavaScript only | Node.js and Python |
| Execution time | Under 1ms | Up to 5-10 seconds |
| Memory | 2 MB | 128 MB to 10 GB |
| Network access | No | Yes |
| Triggers | Viewer Request and Viewer Response only | All 4 triggers |
| Scale | Millions req/sec | Thousands req/sec |
| Cost | Very cheap, has free tier | More expensive, no free tier |
Use CloudFront Functions for: Rewriting URLs or headers at the edge Cache key normalisation (remove irrelevant query strings) Simple JWT token format validation before passing to origin HTTP redirects Use Lambda@Edge for: Complex authentication requiring a network call to a database Reading the request body (POST requests) Calling an AWS service (DynamoDB, Secrets Manager) at the edge Generating dynamic personalised HTML at edge locationsHands-on Lab — Deploy Lambda, Add S3 Trigger, Test VPC Access
Step 1 — Create Lambda execution role
IAM → Roles → Create roleTrusted entity: AWS service → LambdaPermissions: AWSLambdaBasicExecutionRole (allows writing logs to CloudWatch)Role name: devops-lambda-role → CreateStep 2 — Create your first Lambda function
Lambda → Functions → Create functionAuthor from scratchFunction name: devops-order-processorRuntime: Python 3.12Execution role: Use existing → devops-lambda-roleCreate functionStep 3 — Write the function code
In the code editor, replace the default with:import json ## This runs once on cold start — reused across warm invocationsprint("Cold start — initialising") def lambda_handler(event, context): print(f"Received event: {json.dumps(event)}") ## Simulate processing order_id = event.get("orderId", "unknown") city = event.get("city", "unknown") print(f"Processing order {order_id} from {city}") return { "statusCode": 200, "body": json.dumps({ "message": f"Order {order_id} processed", "city": city }) }Click: DeployStep 4 — Test the function
Test → Create new test eventEvent name: test-orderEvent JSON:{"orderId": "ORD-2024-001", "city": "Mumbai", "amount": 450} Click: Test You will see:Response: {"statusCode": 200, "body": "..."}Logs: both "Cold start" and "Processing order" lines appear Click Test again — notice "Cold start" does NOT appear this time.The same environment is reused (warm invocation).Step 5 — Add S3 trigger
Lambda → devops-order-processor → Configuration → Triggers → Add triggerSource: S3Bucket: your existing bucketEvent type: All object create eventsAdd Upload any file to that S3 bucket:echo "test order data" > order.txtaws s3 cp order.txt s3://your-bucket-name/Lambda → devops-order-processor → Monitor → View CloudWatch logsYou should see a new log entry triggered by the S3 upload.Step 6 — Set reserved concurrency
Lambda → devops-order-processor → Configuration → Concurrency → EditReserved concurrency: 100Save This guarantees 100 slots always available for this functionand prevents it from consuming more than 100 slots during a spike.Step 7 — Check metrics in CloudWatch
Lambda → devops-order-processor → MonitorYou can see:Invocations → total times function ranDuration → how long each run tookErrors → any failuresThrottles → any times concurrency limit was hitInit Duration → cold start times (filter for these to see cold vs warm)Step 8 — Cleanup
Lambda → devops-order-processor → Actions → Delete functionIAM → Roles → devops-lambda-role → DeleteS3 → remove trigger from bucket settingsCommon Mistakes to Avoid
Common MistakeCreating database connections inside the handler function. Every invocation — cold and warm — creates a new connection if the code is inside the handler. Move all initialisation (DB connections, SDK clients, config loading) outside the handler. It runs once on cold start and is reused on all subsequent warm invocations at zero cost.
Common MistakeNot setting reserved concurrency on critical functions. Without it, a traffic spike on one function can consume all 1,000 account-level concurrent slots and throttle every other function — including your payment processing API. Reserve concurrency for every production function that has an SLA.
Common MistakeConnecting Lambda directly to RDS at scale without RDS Proxy. 10 concurrent Lambda executions = 10 database connections = fine. 500 concurrent = 500 connections = database refuses connections = application-wide failure. Add RDS Proxy before you need it.
TipUse the Lambda Power Tuning open-source tool to find the optimal memory setting for your function. It runs your function at multiple memory configurations in parallel and shows you the cost-performance curve. Many functions are fastest AND cheapest at a memory setting that is not the default 128 MB. The tool finds the sweet spot automatically.