Apache Kafka for Data Engineers
Learn Kafka for data pipelines: topics, partitions, consumer groups, Kafka Connect, CDC with Debezium, Schema Registry, and how Kinesis compares.
What You'll Learn
Understanding Why Streaming Exists
It is 9:40 PM on the biggest sale night of the year at acme-shop. A customer taps "Pay".
Understanding Kafka Core Concepts
Everything in Kafka is built from four ideas: topics, partitions, replicas, and offsets. Get these four clear and the rest of the module is detail.
Building Producers and Consumers
Writing events with a producer A producer is any application that publishes events to a topic. It picks the topic and, usually, a key.
Moving Data with Kafka Connect
Writing a custom producer or consumer for every system you need to connect would leave you maintaining dozens of near-identical programs.
Capturing Changes with Debezium
Why polling for changes fails A common first approach is to query the database every few minutes: "give me rows where updated_at is newer than my...
Keeping Schemas Safe with Schema Registry
The problem schemas solve To Kafka, a message is opaque bytes.
Skills You'll Master
Curriculum Index12 topics
Understanding Why Streaming Exists
It is 9:40 PM on the biggest sale night of the year at acme-shop. A customer taps "Pay".
Understanding Kafka Core Concepts
Everything in Kafka is built from four ideas: topics, partitions, replicas, and offsets.
Building Producers and Consumers
Writing events with a producer A producer is any application that publishes events to a topic.
Moving Data with Kafka Connect
Writing a custom producer or consumer for every system you need to connect would leave you maintaining dozens of...
Capturing Changes with Debezium
Why polling for changes fails A common first approach is to query the database every few minutes: "give me rows where...
Keeping Schemas Safe with Schema Registry
The problem schemas solve To Kafka, a message is opaque bytes.
Choosing Kafka, Amazon MSK, or Kinesis
On AWS you will meet three ways to run an event stream. Self-managed Kafka means you run brokers yourself.
Understanding Delivery Guarantees and Consumer Lag
Delivery semantics in plain terms At-least-once: a message may arrive more than once (after a retry, say) but is never...
Understanding Log Compaction (Good to Know)
Retention by age versus retention by key Normal retention deletes messages after a time period, whatever they contain.
Running the Hands-on Lab
This lab has three parts: a producer and consumer with a lag check, real CDC from PostgreSQL through Debezium, and a...
Understanding What You Built and What Comes Next
What you built You ran a one-broker Kafka cluster in KRaft mode, produced 2,000 acme-shop orders keyed by customer...
Reviewing the Quick Reference and Common Mistakes
Quick reference Common mistakes Running a host script against a broker that advertises only its container name makes...
Career Impact
Roles that use the skills in this module.
Data Engineer
Platform Engineer
Next Modules
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
Not quite. A queue deletes a message once a consumer has handled it. Kafka keeps events in a log for a retention period, so many independent consumer groups can read the same events, and a new consumer can replay old ones.
Start from the parallelism you need. Each partition is read by at most one consumer in a group, so 3 partitions allow at most 3 active consumers per group. More partitions add broker overhead and slower rebalances, so do not maximise blindly.
No. Kafka 4.0 removed ZooKeeper completely and runs in KRaft mode, where brokers manage cluster metadata themselves. Older clusters you meet at work may still use ZooKeeper.
Polling with an updated_at filter misses deletes and collapses rapid changes into the last state. CDC reads the database transaction log, so every insert, update, and delete arrives in commit order without extra load on the tables.
Choose Kinesis when your stack is all AWS and you want almost no cluster operations. Choose Kafka or Amazon MSK when you need its larger ecosystem (Connect, Debezium, Schema Registry), long retention, or portability across clouds.