Event-Driven Architecture on AWS Explained
Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.
Swiggy receives 50,000 orders per minute during the IPL final dinner rush. Each order triggers seven downstream actions: fraud check, restaurant notification, delivery partner assignment, payment processing, inventory update, email confirmation, and analytics write. If the Order Service calls all seven directly and any one of them is slow or unavailable, the user waits. At 50,000 orders per minute, one slow downstream service turns into 50,000 stalled transactions.
Event-driven architecture solves this. The Order Service publishes one event and moves on. Seven subscribers each receive the event and process it independently. If the email service is slow, orders still complete. If analytics is down, payments still work. Each service scales independently based on its own load.
The Core Problem with Synchronous Architecture
Tight coupling is the root cause. When Service A calls Service B directly, A depends on B's availability, B's latency, and B's capacity.
Synchronous (tightly coupled): Order Service -> HTTP call -> Fraud Service (300ms) -> HTTP call -> Restaurant Service (200ms) -> HTTP call -> Email Service (150ms) -> HTTP call -> Analytics Service (400ms)Total: 1,050ms before confirmation, any service can break the chain Asynchronous (event-driven): Order Service -> publish event -> SNS Topic -> done (10ms)Order Service confirms to user immediately Downstream in parallel (invisible to user):Fraud Service <- reads from SQS Queue ARestaurant Svc <- reads from SQS Queue BEmail Service <- reads from SQS Queue CAnalytics <- Kinesis Firehose -> S3Event-driven architecture converts sequential serial processing into parallel independent processing.
SQS: The Work Queue
SQS is the backbone of most event-driven AWS systems. It decouples producers from consumers and absorbs traffic spikes through persistence.
At Razorpay, webhook events from payment gateways arrive in bursts during peak hours. The processing system reads from SQS at a controlled rate, preventing the database from being overwhelmed during spikes.
Standard queue with a dead letter queue:
## Create a standard SQS queue with DLQaws sqs create-queue \ --queue-name razorpay-webhook-dlq \ --region ap-south-1 aws sqs create-queue \ --queue-name razorpay-webhooks \ --attributes VisibilityTimeout=60,ReceiveMessageWaitTimeSeconds=20 \ --region ap-south-1FIFO queue for ordered processing:
When processing financial transactions, order matters. A sell before a buy on the same account is a compliance violation. SQS FIFO guarantees strict ordering per Message Group ID.
## Create FIFO queue for trade order processingaws sqs create-queue \ --queue-name zerodha-trade-orders.fifo \ --attributes FifoQueue=true,ContentBasedDeduplication=true \ --region ap-south-1Visibility timeout, sized correctly:
Set visibility timeout = 6 x Lambda timeoutLambda timeout: 30 secondsVisibility timeout: 180 seconds minimumSNS: Fan-Out to Multiple Services
SNS delivers one message to all subscribers simultaneously. When an order is placed, every service that cares about the event receives it at the same time.
SNS alone loses messages if a subscriber is down. The production pattern is SNS plus SQS, where SNS fans out to durable SQS queues:
## Create SNS topic for order eventsaws sns create-topic \ --name swiggy-order-events \ --region ap-south-1 ## Create SQS queues for each downstream serviceaws sqs create-queue --queue-name swiggy-fraud-queue --region ap-south-1aws sqs create-queue --queue-name swiggy-restaurant-queue --region ap-south-1 ## Subscribe each queue to the topicaws sns subscribe \ --topic-arn arn:aws:sns:ap-south-1:123456789012:swiggy-order-events \ --protocol sqs \ --notification-endpoint arn:aws:sqs:ap-south-1:123456789012:swiggy-fraud-queue \ --region ap-south-1Message filtering, so each service gets only what it needs:
## Fraud service only processes high-value ordersaws sns set-subscription-attributes \ --subscription-arn arn:aws:sns:ap-south-1:123456789012:fraud-queue-subscription \ --attribute-name FilterPolicy \ --attribute-value fraud-filter.json \ --region ap-south-1The fraud queue only receives orders above Rs. 2,000. Lower-value orders go directly to fulfilment without fraud review. One topic, intelligent routing.
EventBridge: React to Everything That Happens in AWS
EventBridge watches every AWS service event and routes them to targets based on rules you define. No polling. No custom integrations. Near-instant reaction to infrastructure events.
At Zerodha, when a production EC2 instance terminates unexpectedly, an incident is created automatically:
## React to unexpected EC2 terminationaws events put-rule \ --name detect-ec2-termination \ --event-pattern file://ec2-termination-pattern.json \ --state ENABLED \ --region ap-south-1 ## Route to Lambda for incident creationaws events put-targets \ --rule detect-ec2-termination \ --targets file://incident-lambda-target.json \ --region ap-south-1Scheduled triggers, no cron servers:
## Run nightly data cleanup at 2 AM IST (8:30 PM UTC)aws events put-rule \ --name nightly-cleanup \ --schedule-expression "cron(30 20 * * ? *)" \ --state ENABLED \ --region ap-south-1Custom application events:
## Publish custom business events from your applicationaws events put-events \ --entries file://payment-succeeded-event.json \ --region ap-south-1Kinesis: Real-Time Streaming for High Volume
For high-volume event streams — clickstream data, IoT sensor readings, financial tick data — Kinesis Data Streams handles what SQS cannot: multiple consumers reading the same stream simultaneously with replay capability.
## Create Kinesis stream for Hotstar clickstreamaws kinesis create-stream \ --stream-name hotstar-clickstream \ --shard-count 10 \ --region ap-south-110 shards:Ingest: 10 MB/s or 10,000 records/secondRead: 20 MB/s (2 MB/s per shard)Multiple Lambda, Flink, and Firehose consumers read simultaneouslyKinesis Data Firehose delivers the stream to S3 for archiving, no consumer code needed.
Dead Letter Queues: Handle Failures Gracefully
Every event-driven system needs a failure path. Messages that fail processing repeatedly go to a Dead Letter Queue for inspection and manual reprocessing.
Message arrives in SQS queue |Lambda consumer picks it up |Processing fails (downstream service down, parsing error, etc.) |Message returns to queue (visibility timeout expires) |Retried 3 times (maxReceiveCount = 3) |Moved to DLQ automatically |CloudWatch alarm on DLQ depth -> SNS alert -> engineer investigatesSQS Plus ASG: A Self-Regulating Pipeline
The most powerful pattern for batch and background processing: SQS queue depth drives Auto Scaling Group size automatically.
Orders land in SQS queue |CloudWatch monitors ApproximateNumberOfMessages |Queue depth > 1,000 -> alarm fires -> ASG adds 5 EC2 workersQueue depth < 100 -> alarm fires -> ASG removes workers to minimum Result: worker fleet automatically matches workload volume3 AM quiet -> 2 workers runningDinner rush -> 20 workers runningTrade-offs and Alternatives
| Pattern | Persistence | Best For |
|---|---|---|
| SQS | Yes, until consumed | Work queues, one consumer per message |
| SNS + SQS | Yes (via SQS) | Fan-out notifications to many services |
| EventBridge | No (archive optional) | AWS infrastructure event reactions |
| Kinesis | Yes, 1-365 days | High-volume streams, multiple consumers, replay |
Production Implementation Guidelines
- Start with SQS + Lambda for any new background processing job. Lowest-risk, easiest-to-operate combination.
- Use SNS fan-out to SQS when multiple services need the same event. Never let services call each other directly for async events.
- Enable long polling (
ReceiveMessageWaitTimeSeconds=20) on every SQS queue. Short polling wastes API calls and adds latency. - Set Dead Letter Queue
maxReceiveCountto 3. Alert on DLQ depth going above zero. - Use EventBridge for AWS infrastructure events instead of polling CloudTrail with Lambda.
- Design every Lambda consumer to be idempotent. Standard SQS delivers at-least-once, so the same message will be processed more than once eventually.
NoteReferences and Further Reading
- Amazon SQS developer guide — visibility timeout, FIFO, and DLQ patterns
- EventBridge patterns — event filtering and routing reference
- Kinesis Data Streams vs SQS — AWS's own comparison
Frequently Asked Questions
Why does SNS alone lose messages if a subscriber service is down?
SNS delivers a message at the moment of publish with no built-in persistence — if a subscriber isn't available to receive it right then, the message doesn't wait. The production fix is to fan out from SNS into durable SQS queues, so each subscriber has its own persistent buffer to consume from when it recovers.
How should you size the SQS visibility timeout relative to your Lambda consumer's timeout?
Set the visibility timeout to roughly 6x the Lambda function's timeout — for a 30-second Lambda, a 180-second visibility timeout minimum — so a message isn't accidentally redelivered to a second consumer while the first one is still legitimately processing it.
When should you use SQS FIFO instead of a standard queue?
When processing order matters for correctness, not just convenience — financial transactions where a sell processed before a corresponding buy would be a compliance violation are the clearest case. FIFO guarantees strict ordering per Message Group ID, at lower throughput than standard queues.
What's the difference between using EventBridge and Kinesis for high-volume events?
EventBridge is built for routing discrete AWS infrastructure and application events to targets based on rules, with no built-in replay. Kinesis Data Streams is built for high-volume, ordered streaming data — like clickstreams or tick data — where multiple consumers need to read the same stream simultaneously and replay capability matters.
Why must Lambda consumers reading from SQS be idempotent?
Standard SQS guarantees at-least-once delivery, not exactly-once — the same message can be redelivered and processed more than once under normal operation, so a consumer that isn't idempotent risks duplicate side effects like double-charging a payment or double-sending a notification.
Discussion0