What you will learn
- Why decoupling exists — the difference between synchronous and asynchronous communication
- SQS Standard vs FIFO queues — what each guarantees and what each trades away
- Visibility timeout — what it is and the three problems it causes when wrong
- Long Polling vs Short Polling — why you always want long polling
- Dead Letter Queues — what happens to messages that keep failing
- How SQS with Auto Scaling Groups creates a self-regulating pipeline
- SNS pub/sub — the fan-out pattern and how it combines with SQS for durability
- Kinesis Data Streams — shards, retention, and the replay capability SQS does not have
- Amazon Data Firehose — near real-time delivery without writing any consumer code
- SQS vs SNS vs Kinesis — how to choose the right service for the right problem
- Amazon MQ — when open protocol compatibility matters
Why this matters
At Swiggy, when a user places an order, the Order Service needs to notify five downstream services — Fraud Detection, Shipping, Email Notifications, Analytics, and the Restaurant Display. If Order Service calls all five directly and one of them is slow or down, the user waits or gets an error. At dinner rush, when 50,000 orders hit simultaneously, this falls apart completely.
At Zerodha, trade confirmation events must arrive in the exact sequence they were placed — a sell before a buy on the same account is a compliance violation.
At Hotstar, a live cricket match generates 500,000 data points per second that must be captured without any loss for real-time analytics.
These three problems — decoupling, ordering, and real-time streaming — are what SQS, SNS, and Kinesis are each built to solve.
Why Decoupling Exists
When services communicate synchronously, they depend on each other's availability and speed.
Synchronous (tightly coupled): Order Service ─────────────────────────────→ Email Service If Email Service is slow → Order Service waitsIf Email Service is down → Order Service failsTraffic spike → both services overwhelmed at the same time Asynchronous (decoupled): Order Service ──→ [Queue] ──→ Email Service Order Service drops message into queue and moves on immediatelyEmail Service reads from queue at its own paceBoth sides scale independentlyNeither waits for the other to be availableThis is the core idea. A queue, a topic, or a stream sits in the middle. Producers and consumers never talk directly. Each side can be slow, fast, scaled up, or temporarily down without affecting the other.
Amazon SQS — Simple Queue Service
A queue holds messages until consumers read and process them. Producers put messages in. Consumers pull them out.
Producers SQS Queue Consumers[Order Service] ──────→ ||||||||||||||| ──────→ [Payment Service][Order Service] ──────→ ||||||||||||||| ──────→ [Payment Service] messages waiting multiple consumersSQS Standard Queue:
| Property | Value |
|---|---|
| Throughput | Unlimited — no cap on messages per second |
| Message retention | 4 days default, up to 14 days |
| Max message size | 256 KB |
| Delivery guarantee | At-least-once — a message might be delivered more than once |
| Ordering | Best-effort — messages may arrive out of order |
The at-least-once delivery and best-effort ordering are intentional trade-offs for unlimited throughput. Your consumer must be idempotent — processing the same message twice should produce the same result.
SQS FIFO Queue — when order and exactly-once matter:
Standard Queue: Producer sends: Order-1, Order-2, Order-3 Consumer gets: Order-3, Order-1, Order-2 (any order) FIFO Queue: Producer sends: Order-1, Order-2, Order-3 Consumer gets: Order-1, Order-2, Order-3 (guaranteed order)FIFO also prevents duplicates using a Message Deduplication ID. Send the same message ID twice within 5 minutes and the second one is silently dropped.
Throughput limits (the trade-off for ordering):
| Standard | FIFO | |
|---|---|---|
| Throughput | Unlimited | 300 msgs/sec, or 3,000 with batching |
| Order | Best-effort | Strictly guaranteed |
| Duplicates | Can happen | Removed automatically |
| Use for | High volume, tolerates duplicates | Financial transactions, inventory, ordered events |
Visibility Timeout — The Hidden Mechanism
When Consumer A reads a message, the message does not disappear. It becomes invisible to other consumers for a set time. This is the visibility timeout.
Consumer A polls a message ↓Message becomes INVISIBLE to all other consumers ↓Consumer A has 30 seconds (default) to process and delete it ↓Consumer A calls DeleteMessage in time → message gone permanentlyConsumer A fails or times out → message reappears → another consumer picks it upDefault timeout: 30 seconds. Range: 0 seconds to 12 hours.
Three problems caused by wrong timeout:
Too short: Your function needs 45 seconds but timeout is 30 Message reappears before processing finishes Another consumer processes the same message → duplicate processing Fix: set timeout longer than your worst-case processing time Too long: Consumer crashes midway through Message stays invisible for hours → nothing can retry it Fix: do not set unnecessarily high Processing takes longer than expected: You started processing but realise it will take longer Fix: call ChangeMessageVisibility API to extend the timeout before it expiresRememberWhen Lambda is your SQS consumer, set the SQS visibility timeout to at least 6 times the Lambda timeout. Lambda needs time to retry failed batches. Too short a timeout causes the same messages to be processed multiple times simultaneously.
Long Polling — Always Use This
By default, when your consumer calls ReceiveMessage and the queue is empty, SQS returns immediately with no messages. This is short polling.
Short Polling: Consumer asks "any messages?" → SQS: "no" → Consumer asks again Repeats every second → wasted API calls → cost adds up → slower response Long Polling: Consumer asks "any messages, I'll wait up to 20 seconds" Queue empty → SQS holds the connection open Message arrives → SQS returns it immediately No message in 20 seconds → returns empty onceLong Polling is cheaper, faster, and the right default. Set it at the queue level (Receive message wait time: 20 seconds) or per-API-call (WaitTimeSeconds=20).
Dead Letter Queue — What Happens to Failed Messages
Some messages will always fail. A malformed payload. A bug in your consumer. A downstream service permanently unavailable. Without a DLQ, these messages loop in your queue forever — consuming resources and blocking other messages.
A Dead Letter Queue receives messages that have been received too many times without being deleted.
Normal Queue (maxReceiveCount = 3): Message received → processing fails → message returns to queue Received again → fails again → returns again Received a third time → fails again → MOVED TO DLQ automatically ↓DLQ holds the failed messageYou inspect it, find the bug, fix the consumerYou redrive the messages from DLQ back to the normal queue Setting up a DLQ: SQS → your queue → Dead-letter queue → Edit Select your DLQ → set maxReceiveCount to 3 → SaveTipSet up a CloudWatch alarm on the DLQ's ApproximateNumberOfMessagesVisible metric. When messages land in the DLQ, you want to know immediately — not discover it days later when a backlog of failed messages has built up.
SQS with Auto Scaling Groups — The Self-Regulating Pipeline
The most powerful SQS pattern for worker applications.
SQS Queue receives messages (orders, jobs, events) ↓CloudWatch monitors: ApproximateNumberOfMessages ↓Queue growing → CloudWatch alarm fires → ASG adds more worker EC2sQueue shrinking → CloudWatch alarm fires → ASG removes worker EC2s ↓Workers always match the volume of work available Result: Queue deep at 9 AM → 20 workers running Queue empty at 3 AM → 2 workers running (minimum) Cost automatically matches workloadNo manual scaling. No fixed fleet running overnight. Infrastructure matches work volume automatically.
SQS as a Buffer to Protect Databases
A database can handle N writes per second. Traffic spikes can easily send 10N writes per second. Without buffering, the database gets overwhelmed and writes fail.
Without SQS: Traffic spike → 10,000 writes/second hit RDS RDS overwhelmed → writes fail → data lost With SQS as buffer: Traffic spike → 10,000 writes/second go into SQS queue (queue absorbs spike) Consumer reads at a safe 500/second → inserts steadily into RDS No data lost. RDS never overwhelmed. Queue slowly drains.TipAny time you hear "too many requests flooding a database or service during a spike" in an architecture question, think SQS as buffer. The queue absorbs the spike and the consumer processes at a safe rate.
Amazon SNS — Simple Notification Service
SQS is a queue — each message goes to one consumer. SNS is pub/sub — one message goes to all subscribers simultaneously.
Without SNS — Order Service calls everyone directly: Order Service → Fraud Detection Order Service → Shipping Service Order Service → Email Service Order Service → Analytics Order Service → SQS Queue Adding a new service requires changing Order Service code every time With SNS — Order Service publishes once: Order Service → SNS Topic → Fraud Detection → Shipping Service → Email Service → Analytics → SQS Queue Order Service only knows about SNS Adding a new subscriber never touches Order Service codeSNS key facts:
Up to 12,500,000 subscriptions per topicUp to 100,000 topics per accountMessages are NOT persisted — if delivery fails, the message is goneAll subscribers receive all messages unless filtering is appliedSNS subscribers:
SQS Queues, Lambda functions, Kinesis Data FirehoseEmail, SMS, mobile push (Apple APNS, Google FCM)HTTP/HTTPS endpointsSNS Fan-Out Pattern — SNS + SQS Together
SNS broadcasts and moves on. If a subscriber is temporarily down when SNS sends, that message is lost forever.
Solution: combine SNS and SQS. SNS fans out to multiple SQS queues. Each queue persists the message for its consumer.
Order placed → SNS Topic publishes once ↓SQS Queue A → Fraud Detection worker (message persists until processed)SQS Queue B → Shipping worker (message persists until processed)SQS Queue C → Email worker (message persists until processed) Benefits: Push once to SNS — fans out automatically to all queues Each queue persists the message independently If Email worker is down, its queue holds the message until it recovers Add a new queue without changing any existing codeS3 event fan-out — the classic use case:
S3 only allows one destination per event rule. You cannot point one S3 event directly at three different Lambda functions.
S3 object uploaded ↓SNS Topic (S3 sends here) ↓SQS Queue 1 → Lambda: resize imageSQS Queue 2 → Lambda: virus scanSQS Queue 3 → Lambda: backup to cold storageSNS sends to one place. Everything else fans out from there.
SNS Message Filtering:
By default all subscribers receive all messages. A filter policy on the subscription tells SNS to only deliver messages that match specific attributes.
SNS Topic (all order events) ├── Filter: status=placed → SQS Fraud Queue (placed orders only) ├── Filter: status=cancelled → SQS Refund Queue (cancellations only) └── No filter → SQS Analytics Queue (receives everything)One SNS topic handles all event types. Each subscriber gets only what it cares about.
RememberThe SQS queue's access policy must allow SNS to write to it. Without this permission, SNS cannot deliver messages. This is the most commonly forgotten step when setting up the fan-out pattern — everything looks configured but messages never arrive in the queue.
Amazon Kinesis Data Streams — Real-Time Streaming
SQS and SNS are for message passing between services. Kinesis is for collecting, storing, and processing continuous streams of data in real time — clicks, sensor readings, application logs, financial transactions, game events.
What is a shard:
A shard is the capacity unit of a Kinesis stream. You provision a certain number of shards and that determines your throughput.
Each shard: Ingest: up to 1 MB/s or 1,000 records/second Read: up to 2 MB/s Need to handle 5 MB/s incoming → provision 5 shardsKey capabilities:
Retention: 1 to 365 days — data stays and can be reprocessedReplay: re-read data from any point in the retention window → processing had a bug? fix it and replay all of yesterday's data → SQS cannot replay — once consumed, gone foreverOrdering: records with same Partition Key always go to same shard, in orderImmutable: data cannot be deleted manually, only expiresTwo capacity modes:
Provisioned Mode: You choose number of shards Pay per shard per hour Scale by adding or removing shards manually Good for: predictable, steady traffic On-Demand Mode: No shard management Auto-scales based on peak traffic over last 30 days Pay per GB in and out Good for: unpredictable or variable trafficRememberKinesis Data Streams is the only messaging service on AWS that supports replay. SQS deletes messages once consumed. SNS never stores them. Kinesis keeps them for up to 365 days and lets you re-read from any point. This makes it the right choice whenever "what if we need to reprocess data" is a requirement.
Amazon Data Firehose — Delivery Without Consumer Code
Kinesis Data Streams requires you to write consumer code — an application that reads from the stream and does something with the data. Firehose removes this entirely.
Firehose takes streaming data and loads it into a destination automatically. No consumers. No servers. Fully managed.
Producers send data to Firehose ↓Firehose buffers by size or time (e.g. 5 MB or 60 seconds, whichever comes first) ↓Optional Lambda function transforms each record ↓Delivers batch to destinationDestinations:
AWS: S3, Redshift (via S3 COPY), OpenSearchThird-party: Splunk, Datadog, MongoDB, New RelicCustom: Any HTTP endpointNear real-time — not real-time:
Firehose buffers before delivering. The minimum buffer is 60 seconds. This is why it is "near real-time" — not instant. For millisecond-latency streaming, use Kinesis Data Streams with your own consumer.
Kinesis Data Streams vs Firehose:
| Kinesis Data Streams | Firehose | |
|---|---|---|
| Purpose | Collect and store streaming data | Deliver streaming data to a destination |
| Consumer code | You write it | Not needed — fully managed |
| Latency | Milliseconds | 60+ seconds (buffered) |
| Data storage | 1 to 365 days | No storage — routes and delivers |
| Replay | Yes | No |
| Best for | Real-time processing, replay needed | Archiving to S3, loading into Redshift |
They work great together. Kinesis Data Streams collects and holds the data with replay capability. Firehose reads from the stream and delivers to S3 or Redshift. You get both replay and managed delivery.
SQS vs SNS vs Kinesis — The Decision Guide
| SQS | SNS | Kinesis | |
|---|---|---|---|
| Model | Queue — pull | Pub/Sub — push | Stream — pull |
| Message persistence | Until consumer deletes | Not persisted | 1 to 365 days |
| Each message goes to | One consumer | All subscribers | Multiple consumers per shard |
| Replay | No | No | Yes |
| Throughput | Unlimited | Unlimited | Per shard (1 MB/s in) |
Choose SQS when: Work queue — one consumer processes each job Messages must persist until processed Need to buffer a spike against a downstream service Choose SNS when: Same event needs to reach multiple services at once Fan-out to many receivers Push notifications to mobile or email Choose Kinesis when: Continuous high-volume streaming data (clicks, IoT, logs, metrics) You need to replay data from the past Multiple consumers need the same stream simultaneously Real-time analytics on data as it arrivesAmazon MQ — For Open Protocol Migration
SQS and SNS use AWS-proprietary APIs. Old on-premises applications often use open messaging protocols — MQTT, AMQP, STOMP, OpenWire. Migrating such an application to use SQS would mean rewriting the messaging layer completely.
Amazon MQ provides managed RabbitMQ and ActiveMQ brokers that speak those open protocols natively. Zero code changes needed.
Old on-premises app using AMQP → Amazon MQ (same protocol) → zero code changesNew cloud-native app → SQS/SNS → purpose-built, massively scalable| Amazon MQ | SQS and SNS | |
|---|---|---|
| Protocols | MQTT, AMQP, STOMP, OpenWire, WSS | AWS APIs only |
| Best for | Migrating existing on-prem apps | New cloud-native applications |
| Scale | Limited | Virtually unlimited |
Amazon MQ High Availability:
AZ-1a: Active Broker ← clients connect hereAZ-1b: Standby Broker (synchronised via EFS shared storage) Active fails → standby takes over automaticallyBoth share the same EFS storage → no messages lost on failoverRememberUse Amazon MQ only when migrating from an existing on-premises application that uses open protocols and you cannot change the messaging layer. For any new application built from scratch on AWS, use SQS and SNS — they scale massively and require no server management.
Hands-on Lab — SQS Queue, DLQ, SNS Fan-Out
Step 1 — Create a Dead Letter Queue
SQS → Create queueType: StandardName: devops-orders-dlqAll other settings: defaultCreate queue Copy the DLQ ARN — you will need it in the next step.Step 2 — Create the main queue with DLQ configured
SQS → Create queueType: StandardName: devops-orders Scroll to Dead-letter queue:Enable → select devops-orders-dlqMaximum receives: 3 → Save Scroll to Visibility timeout: set to 60 secondsReceive message wait time: 20 seconds (enables long polling)Create queueStep 3 — Send a test message
SQS → Queues → devops-orders → Send and receive messagesMessage body: {"orderId": "ORD-001", "city": "Mumbai", "amount": 450}Send messageStep 4 — Receive and inspect the message
Same screen → Poll for messages → Start polling for messagesMessage appears → click to view → read the bodyDo NOT delete it yet — watch what happens to visibility timeout Wait 60 seconds → the message reappears (visibility timeout expired)This simulates a consumer that failed to process and deleteStep 5 — Create an SNS Topic and fan-out to two queues
SNS → Topics → Create topicType: StandardName: devops-order-eventsCreate topic Create a second SQS queue: devops-fraud-check (same settings as devops-orders) SNS → devops-order-events → Create subscriptionProtocol: SQSEndpoint: ARN of devops-orders queue → Create subscription Create another subscription:Protocol: SQSEndpoint: ARN of devops-fraud-check queue → Create subscriptionStep 6 — Publish to SNS and verify fan-out
SNS → devops-order-events → Publish messageMessage body: {"orderId": "ORD-002", "event": "placed"}Publish message Check devops-orders queue → Poll for messages → message should appearCheck devops-fraud-check queue → Poll for messages → same message should appear One publish → both queues received a copy. Fan-out confirmed.Step 7 — Cleanup
SQS → Queues → select all three → DeleteSNS → Topics → devops-order-events → DeleteSubscriptions are deleted automatically with the topicCommon Mistakes to Avoid
Common MistakeChoosing SQS when multiple services need the same message. SQS sends each message to only one consumer. If Fraud Detection, Shipping, and Email all need the same order event, you need SNS to fan-out to three separate SQS queues — one per service. Each service gets its own copy of the message with its own processing lifecycle.
Common MistakeNot calling DeleteMessage after processing in SQS. If your consumer processes the message successfully but forgets to delete it, SQS makes it visible again after the visibility timeout. Another consumer picks it up and processes it a second time. Always delete after successful processing. Always.
Common MistakeUsing Kinesis Firehose when you need Kinesis Data Streams. Firehose buffers for at least 60 seconds and delivers to a destination — it does not store data and cannot replay. If you need real-time processing with millisecond latency or the ability to re-read past data, you need Kinesis Data Streams with your own consumer. Firehose is delivery, not streaming.
TipFor a real-time data pipeline — IoT device data, application clickstreams, live metrics — the full serverless architecture is: devices send to Kinesis Data Streams → Firehose reads from the stream and delivers to S3 in Parquet format → Athena queries the Parquet files for analytics → QuickSight builds dashboards. No servers anywhere in the pipeline. Kinesis handles the real-time collection and replay capability. Firehose handles the format conversion and S3 delivery.