Skip to main content

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

Swiggy receives 50,000 orders per minute during the IPL final dinner rush. Each order triggers seven downstream actions: fraud check, restaurant notification, delivery partner assignment, payment processing, inventory update, email confirmation, and analytics write. If the Order Service calls all seven directly and any one of them is slow or unavailable, the user waits. At 50,000 orders per minute, one slow downstream service turns into 50,000 stalled transactions.

Event-driven architecture solves this. The Order Service publishes one event and moves on. Seven subscribers each receive the event and process it independently. If the email service is slow, orders still complete. If analytics is down, payments still work. Each service scales independently based on its own load.

The Core Problem with Synchronous Architecture

Tight coupling is the root cause. When Service A calls Service B directly, A depends on B's availability, B's latency, and B's capacity.

◈ DIAGRAM
Synchronous (tightly coupled):
Order Service -> HTTP call -> Fraud Service (300ms)
-> HTTP call -> Restaurant Service (200ms)
-> HTTP call -> Email Service (150ms)
-> HTTP call -> Analytics Service (400ms)
Total: 1,050ms before confirmation, any service can break the chain
Asynchronous (event-driven):
Order Service -> publish event -> SNS Topic -> done (10ms)
Order Service confirms to user immediately
Downstream in parallel (invisible to user):
Fraud Service <- reads from SQS Queue A
Restaurant Svc <- reads from SQS Queue B
Email Service <- reads from SQS Queue C
Analytics <- Kinesis Firehose -> S3

Event-driven architecture converts sequential serial processing into parallel independent processing.

SQS: The Work Queue

SQS is the backbone of most event-driven AWS systems. It decouples producers from consumers and absorbs traffic spikes through persistence.

At Razorpay, webhook events from payment gateways arrive in bursts during peak hours. The processing system reads from SQS at a controlled rate, preventing the database from being overwhelmed during spikes.

Standard queue with a dead letter queue:

Bash
## Create a standard SQS queue with DLQ
aws sqs create-queue \
--queue-name razorpay-webhook-dlq \
--region ap-south-1
aws sqs create-queue \
--queue-name razorpay-webhooks \
--attributes VisibilityTimeout=60,ReceiveMessageWaitTimeSeconds=20 \
--region ap-south-1

FIFO queue for ordered processing:

When processing financial transactions, order matters. A sell before a buy on the same account is a compliance violation. SQS FIFO guarantees strict ordering per Message Group ID.

Bash
## Create FIFO queue for trade order processing
aws sqs create-queue \
--queue-name zerodha-trade-orders.fifo \
--attributes FifoQueue=true,ContentBasedDeduplication=true \
--region ap-south-1

Visibility timeout, sized correctly:

TEXT
Set visibility timeout = 6 x Lambda timeout
Lambda timeout: 30 seconds
Visibility timeout: 180 seconds minimum

SNS: Fan-Out to Multiple Services

SNS delivers one message to all subscribers simultaneously. When an order is placed, every service that cares about the event receives it at the same time.

SNS alone loses messages if a subscriber is down. The production pattern is SNS plus SQS, where SNS fans out to durable SQS queues:

Bash
## Create SNS topic for order events
aws sns create-topic \
--name swiggy-order-events \
--region ap-south-1
## Create SQS queues for each downstream service
aws sqs create-queue --queue-name swiggy-fraud-queue --region ap-south-1
aws sqs create-queue --queue-name swiggy-restaurant-queue --region ap-south-1
## Subscribe each queue to the topic
aws sns subscribe \
--topic-arn arn:aws:sns:ap-south-1:123456789012:swiggy-order-events \
--protocol sqs \
--notification-endpoint arn:aws:sqs:ap-south-1:123456789012:swiggy-fraud-queue \
--region ap-south-1

Message filtering, so each service gets only what it needs:

Bash
## Fraud service only processes high-value orders
aws sns set-subscription-attributes \
--subscription-arn arn:aws:sns:ap-south-1:123456789012:fraud-queue-subscription \
--attribute-name FilterPolicy \
--attribute-value fraud-filter.json \
--region ap-south-1

The fraud queue only receives orders above Rs. 2,000. Lower-value orders go directly to fulfilment without fraud review. One topic, intelligent routing.

EventBridge: React to Everything That Happens in AWS

EventBridge watches every AWS service event and routes them to targets based on rules you define. No polling. No custom integrations. Near-instant reaction to infrastructure events.

At Zerodha, when a production EC2 instance terminates unexpectedly, an incident is created automatically:

Bash
## React to unexpected EC2 termination
aws events put-rule \
--name detect-ec2-termination \
--event-pattern file://ec2-termination-pattern.json \
--state ENABLED \
--region ap-south-1
## Route to Lambda for incident creation
aws events put-targets \
--rule detect-ec2-termination \
--targets file://incident-lambda-target.json \
--region ap-south-1

Scheduled triggers, no cron servers:

Bash
## Run nightly data cleanup at 2 AM IST (8:30 PM UTC)
aws events put-rule \
--name nightly-cleanup \
--schedule-expression "cron(30 20 * * ? *)" \
--state ENABLED \
--region ap-south-1

Custom application events:

Bash
## Publish custom business events from your application
aws events put-events \
--entries file://payment-succeeded-event.json \
--region ap-south-1

Kinesis: Real-Time Streaming for High Volume

For high-volume event streams — clickstream data, IoT sensor readings, financial tick data — Kinesis Data Streams handles what SQS cannot: multiple consumers reading the same stream simultaneously with replay capability.

Bash
## Create Kinesis stream for Hotstar clickstream
aws kinesis create-stream \
--stream-name hotstar-clickstream \
--shard-count 10 \
--region ap-south-1
TEXT
10 shards:
Ingest: 10 MB/s or 10,000 records/second
Read: 20 MB/s (2 MB/s per shard)
Multiple Lambda, Flink, and Firehose consumers read simultaneously

Kinesis Data Firehose delivers the stream to S3 for archiving, no consumer code needed.

Dead Letter Queues: Handle Failures Gracefully

Every event-driven system needs a failure path. Messages that fail processing repeatedly go to a Dead Letter Queue for inspection and manual reprocessing.

◈ DIAGRAM
Message arrives in SQS queue
|
Lambda consumer picks it up
|
Processing fails (downstream service down, parsing error, etc.)
|
Message returns to queue (visibility timeout expires)
|
Retried 3 times (maxReceiveCount = 3)
|
Moved to DLQ automatically
|
CloudWatch alarm on DLQ depth -> SNS alert -> engineer investigates

SQS Plus ASG: A Self-Regulating Pipeline

The most powerful pattern for batch and background processing: SQS queue depth drives Auto Scaling Group size automatically.

◈ DIAGRAM
Orders land in SQS queue
|
CloudWatch monitors ApproximateNumberOfMessages
|
Queue depth > 1,000 -> alarm fires -> ASG adds 5 EC2 workers
Queue depth < 100 -> alarm fires -> ASG removes workers to minimum
Result: worker fleet automatically matches workload volume
3 AM quiet -> 2 workers running
Dinner rush -> 20 workers running

Trade-offs and Alternatives

Pattern Persistence Best For
SQS Yes, until consumed Work queues, one consumer per message
SNS + SQS Yes (via SQS) Fan-out notifications to many services
EventBridge No (archive optional) AWS infrastructure event reactions
Kinesis Yes, 1-365 days High-volume streams, multiple consumers, replay

Production Implementation Guidelines

  • Start with SQS + Lambda for any new background processing job. Lowest-risk, easiest-to-operate combination.
  • Use SNS fan-out to SQS when multiple services need the same event. Never let services call each other directly for async events.
  • Enable long polling (ReceiveMessageWaitTimeSeconds=20) on every SQS queue. Short polling wastes API calls and adds latency.
  • Set Dead Letter Queue maxReceiveCount to 3. Alert on DLQ depth going above zero.
  • Use EventBridge for AWS infrastructure events instead of polling CloudTrail with Lambda.
  • Design every Lambda consumer to be idempotent. Standard SQS delivers at-least-once, so the same message will be processed more than once eventually.
Note

References and Further Reading

Frequently Asked Questions

Why does SNS alone lose messages if a subscriber service is down?

SNS delivers a message at the moment of publish with no built-in persistence — if a subscriber isn't available to receive it right then, the message doesn't wait. The production fix is to fan out from SNS into durable SQS queues, so each subscriber has its own persistent buffer to consume from when it recovers.

How should you size the SQS visibility timeout relative to your Lambda consumer's timeout?

Set the visibility timeout to roughly 6x the Lambda function's timeout — for a 30-second Lambda, a 180-second visibility timeout minimum — so a message isn't accidentally redelivered to a second consumer while the first one is still legitimately processing it.

When should you use SQS FIFO instead of a standard queue?

When processing order matters for correctness, not just convenience — financial transactions where a sell processed before a corresponding buy would be a compliance violation are the clearest case. FIFO guarantees strict ordering per Message Group ID, at lower throughput than standard queues.

What's the difference between using EventBridge and Kinesis for high-volume events?

EventBridge is built for routing discrete AWS infrastructure and application events to targets based on rules, with no built-in replay. Kinesis Data Streams is built for high-volume, ordered streaming data — like clickstreams or tick data — where multiple consumers need to read the same stream simultaneously and replay capability matters.

Why must Lambda consumers reading from SQS be idempotent?

Standard SQS guarantees at-least-once delivery, not exactly-once — the same message can be redelivered and processed more than once under normal operation, so a consumer that isn't idempotent risks duplicate side effects like double-charging a payment or double-sending a notification.

Discussion0