What you will learn
- Why Amazon MQ exists — the open protocol migration problem SQS cannot solve
- Amazon MQ with RabbitMQ and ActiveMQ — what each supports
- Amazon MQ High Availability — active/standby with EFS shared storage
- Amazon MSK — managed Apache Kafka and when Kafka wins over Kinesis
- MSK Provisioned vs MSK Serverless
- MSK Connect — running Kafka Connect connectors fully managed
- Amazon Managed Flink — real-time stream processing on Kinesis and MSK
- Flink checkpoints vs savepoints
- MSK vs Kinesis Data Streams — the decision guide
- Why Managed Flink cannot read from Kinesis Firehose directly
Why this matters
A bank running on-premises has its entire transaction processing system built on ActiveMQ — hundreds of thousands of lines of code that produce and consume messages using AMQP. Migrating to SQS would mean rewriting the messaging layer entirely. Amazon MQ runs ActiveMQ as a managed service — the bank moves its broker to AWS and zero application code changes. At a Bangalore media company, their analytics pipeline uses Apache Kafka because it was built by engineers who knew Kafka deeply. MSK runs Kafka as a managed service — no Kafka cluster to provision, patch, or scale. Understanding when open protocol compatibility matters versus when to use AWS-native services is what makes architecture decisions correct.
Why Amazon MQ Exists
SQS and SNS are excellent — but they use AWS-proprietary APIs. Legacy applications built on open messaging protocols cannot use SQS without significant code changes.
Open messaging protocols:
MQTT → IoT devices, lightweight publish-subscribeAMQP → enterprise messaging (RabbitMQ, many banks use this)STOMP → simple text-based messagingOpenWire → ActiveMQ's native protocolWSS → WebSocket-based messagingThe migration problem:
On-premises app using AMQP → wants to move to AWSOption 1: Rewrite to use SQS/SNS APIs → months of work, high riskOption 2: Amazon MQ → zero code changes, AMQP still worksAmazon MQ runs managed RabbitMQ or ActiveMQ brokers in AWS. Your existing applications connect to Amazon MQ exactly as they connected to the on-premises broker — same protocol, same client libraries, same API.
RememberUse Amazon MQ only when migrating existing applications that use open protocols and you cannot change the client code. For any new application built from scratch on AWS, use SQS and SNS — they scale massively and require no server management.
Amazon MQ with RabbitMQ
What RabbitMQ provides:
Exchange routing: Direct, Fanout, Topic, HeadersQueues with durability, persistence, and TTLConsumer acknowledgements and dead letter exchangesAMQP 0-9-1 and AMQP 1.0 protocolsManagement UI for queue and consumer visibilityAmazon MQ for RabbitMQ:
Single-instance broker: one broker, one AZ — for developmentCluster deployment: 3 brokers across 3 AZs — for production Cluster deployment:Broker A (ap-south-1a) ←→ Broker B (ap-south-1b) ←→ Broker C (ap-south-1c)Clients connect to any broker — all are active simultaneouslyOne broker fails → remaining two continue servingData replicated across all three brokers via Mirrored QueuesAmazon MQ with ActiveMQ
What ActiveMQ provides:
JMS (Java Message Service) APIOpenWire, STOMP, AMQP, MQTT protocols simultaneouslyVirtual topics, composite destinationsPersistent messaging with KahaDBAmazon MQ for ActiveMQ High Availability:
Active/Standby deployment: Primary Broker (ap-south-1a) ← clients connect here ↓Amazon EFS (shared storage — both brokers see same messages) ↓Standby Broker (ap-south-1b) ← waiting, synchronised via EFS Primary fails → automatic failover to Standby in ~30 secondsClients reconnect using failover transport URLNo message loss — EFS held the messages throughoutRememberAmazon MQ for ActiveMQ uses shared EFS storage. Both the active and standby broker access the same EFS filesystem — so when failover happens, the standby already has all messages. This is the key difference from running ActiveMQ yourself where you must configure replication manually.
Amazon MSK — Managed Streaming for Kafka
Apache Kafka is the industry standard for high-throughput, durable, replayable event streaming. MSK runs Kafka as a fully managed AWS service.
What MSK gives you:
Kafka brokers managed by AWS — no EC2 instances to provisionAutomatic replacement of failed brokersAutomatic storage scalingIntegration with IAM, VPC, CloudWatchMSK Connect for managed Kafka Connect connectorsSame Kafka API your engineers already knowMSK Provisioned:
You choose the broker instance type, number of brokers, and storage. AWS manages everything else.
Broker type: kafka.m5.large (standard production)Number of brokers: 3 (one per AZ for HA)Storage per broker: 1 TBReplication factor: 3 (each partition replicated across all 3 brokers)MSK Serverless:
No broker sizing or capacity planning. MSK scales automatically based on traffic.
No instance type to chooseNo storage to sizePay per GB in and per GB outUse for: unpredictable workloads, development, teams new to Kafka| MSK Provisioned | MSK Serverless | |
|---|---|---|
| Capacity planning | Required | None |
| Cost model | Per broker per hour | Per GB in/out |
| Best for | Stable predictable throughput | Variable or unknown throughput |
MSK Connect — Managed Kafka Connect
Kafka Connect moves data between Kafka and external systems — databases, S3, Elasticsearch, Salesforce — using pre-built connectors.
Without MSK Connect: run Kafka Connect on EC2, manage workers, handle failures yourself.
With MSK Connect: deploy a connector configuration, AWS runs the workers.
S3 Sink Connector: MSK topic → MSK Connect (S3 Sink) → S3 bucket Every Kafka message archived to S3 automatically Debezium Source Connector: PostgreSQL database → MSK Connect (Debezium) → MSK topic Every database change streams to Kafka in real time (CDC)Amazon Managed Flink — Real-Time Stream Processing
Apache Flink is a framework for processing streaming data in real time. Managed Flink runs Flink as a managed AWS service on top of Kinesis Data Streams or MSK.
What Flink can do:
Transform and enrich streaming events in real timeAggregate streams over time windows (sum last 5 minutes of transactions)Join two streams together (match orders to payments in real time)Detect patterns in event sequences (fraud: 3 failed logins then a large transfer)Filter and route events to different destinationsCheckpoints vs Savepoints:
Checkpoints: Automatic, periodic snapshots taken by Flink Used for fault tolerance — if Flink restarts, it resumes from last checkpoint Managed automatically — no manual action Savepoints: Manual snapshots triggered by you Used for: upgrade Flink version, change the job configuration, maintenance The job pauses, snapshot taken, then resumesRememberManaged Flink reads from Kinesis Data Streams and MSK — it does NOT read from Kinesis Data Firehose. Firehose is a delivery service (source to destination), not a stream you can consume. If you need Flink to process the data, send it to Kinesis Data Streams first. Flink reads from Streams, processes, and can output to Firehose for delivery to S3 or Redshift.
MSK vs Kinesis Data Streams — The Decision Guide
Both are real-time streaming services. The choice comes down to the protocol and ecosystem.
| Amazon MSK | Kinesis Data Streams | |
|---|---|---|
| Protocol | Apache Kafka (open standard) | AWS proprietary |
| Portability | Kafka runs anywhere | AWS only |
| Ecosystem | Full Kafka ecosystem (connectors, Flink, etc) | AWS ecosystem |
| Retention | Configurable (days to forever) | 1 to 365 days |
| Ordering | Per partition | Per shard |
| Throughput unit | Partition | Shard |
| Management | More configuration options | Simpler, fewer options |
| Best for | Teams that know Kafka, multi-cloud, open standard needed | AWS-native teams, simpler streaming |
Choose MSK when: Team already knows Kafka and has Kafka expertise Using Kafka Connect connectors (Debezium, S3 Sink, etc) Need portability — application might run on other clouds Need Kafka-specific features (compacted topics, exactly-once semantics) Choose Kinesis Data Streams when: New to streaming, want simplest option AWS-native only, no Kafka expertise in team Already using other Kinesis services (Firehose, Analytics)Hands-on Lab — Amazon MQ and MSK Console Walkthrough
Step 1 — Explore Amazon MQ options
Amazon MQ → Create brokerEngine type: RabbitMQ → review options: Single-instance (dev): one broker, cheaper Cluster (prod): 3 brokers across 3 AZsDo not create — review only Engine type: ActiveMQ → review options: Single-instance (dev) Active/Standby (prod): note the shared storage configurationDo not create — review onlyStep 2 — Create an MSK Serverless cluster
Amazon MSK → Create clusterCreation method: Quick createCluster type: ServerlessCluster name: devops-kafka-serverlessVPC: your VPC Subnets: select 2 private subnetsCreate cluster — takes 10-15 minutesStep 3 — View cluster details
MSK → devops-kafka-serverless → View client informationBootstrap servers: shows the broker endpoints your producers and consumers connect toAuthentication: IAM (MSK Serverless uses IAM authentication only)Step 4 — Create a topic using the console
MSK → devops-kafka-serverless → Topics → Create topicTopic name: devops-ordersReplication factor: 2Number of partitions: 3Create topicStep 5 — Review MSK Connect
MSK → MSK Connect → Create connectorBrowse available connectors — see S3 Sink, Debezium (PostgreSQL, MySQL), etcDo not create — review the connector configuration optionsNote: MSK Connect runs Kafka Connect workers fully managed by AWSStep 6 — Cleanup
MSK → devops-kafka-serverless → Delete clusterConfirm deletion — this terminates the cluster and all topicsCommon Mistakes to Avoid
Common MistakeUsing Amazon MQ for a new application. If you are building a new application from scratch, use SQS and SNS. They scale to virtually unlimited throughput, require no server management, and cost less. Amazon MQ exists specifically for migrating existing applications that use open protocols — not for new development.
Common MistakeTrying to use Managed Flink to read from Kinesis Data Firehose. Firehose is a delivery pipeline — data goes in and flows out to a destination. It is not a stream you can subscribe to or read from. If you need to process the data with Flink before it is delivered, send it to Kinesis Data Streams. Flink reads from Streams. Streams feeds into Firehose if needed for final delivery.
TipFor a common real-time pipeline combining MSK and Flink: applications produce events to MSK topics → Managed Flink reads from MSK → processes and transforms in real time → outputs enriched events back to MSK or directly to S3 via Kinesis Firehose. This is the pattern many large-scale streaming analytics platforms use on AWS.