Skip to main content

Apache Kafka for Data Engineers

Learn Kafka for data pipelines: topics, partitions, consumer groups, Kafka Connect, CDC with Debezium, Schema Registry, and how Kinesis compares.

~4 hours
12 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Streaming Exists

It is 9:40 PM on the biggest sale night of the year at acme-shop. A customer taps "Pay".

Understanding Kafka Core Concepts

Everything in Kafka is built from four ideas: topics, partitions, replicas, and offsets. Get these four clear and the rest of the module is detail.

Building Producers and Consumers

Writing events with a producer A producer is any application that publishes events to a topic. It picks the topic and, usually, a key.

Moving Data with Kafka Connect

Writing a custom producer or consumer for every system you need to connect would leave you maintaining dozens of near-identical programs.

Capturing Changes with Debezium

Why polling for changes fails A common first approach is to query the database every few minutes: "give me rows where updated_at is newer than my...

Keeping Schemas Safe with Schema Registry

The problem schemas solve To Kafka, a message is opaque bytes.

Skills You'll Master

KAFKASTREAMINGCDCDEBEZIUMKINESIS

Curriculum Index12 topics

1

Understanding Why Streaming Exists

It is 9:40 PM on the biggest sale night of the year at acme-shop. A customer taps "Pay".

2

Understanding Kafka Core Concepts

Everything in Kafka is built from four ideas: topics, partitions, replicas, and offsets.

3

Building Producers and Consumers

Writing events with a producer A producer is any application that publishes events to a topic.

4

Moving Data with Kafka Connect

Writing a custom producer or consumer for every system you need to connect would leave you maintaining dozens of...

5

Capturing Changes with Debezium

Why polling for changes fails A common first approach is to query the database every few minutes: "give me rows where...

6

Keeping Schemas Safe with Schema Registry

The problem schemas solve To Kafka, a message is opaque bytes.

7

Choosing Kafka, Amazon MSK, or Kinesis

On AWS you will meet three ways to run an event stream. Self-managed Kafka means you run brokers yourself.

8

Understanding Delivery Guarantees and Consumer Lag

Delivery semantics in plain terms At-least-once: a message may arrive more than once (after a retry, say) but is never...

9

Understanding Log Compaction (Good to Know)

Retention by age versus retention by key Normal retention deletes messages after a time period, whatever they contain.

10

Running the Hands-on Lab

This lab has three parts: a producer and consumer with a lag check, real CDC from PostgreSQL through Debezium, and a...

11

Understanding What You Built and What Comes Next

What you built You ran a one-broker Kafka cluster in KRaft mode, produced 2,000 acme-shop orders keyed by customer...

12

Reviewing the Quick Reference and Common Mistakes

Quick reference Common mistakes Running a host script against a broker that advertises only its container name makes...

Career Impact

Roles that use the skills in this module.

  • Data Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Not quite. A queue deletes a message once a consumer has handled it. Kafka keeps events in a log for a retention period, so many independent consumer groups can read the same events, and a new consumer can replay old ones.

Start from the parallelism you need. Each partition is read by at most one consumer in a group, so 3 partitions allow at most 3 active consumers per group. More partitions add broker overhead and slower rebalances, so do not maximise blindly.

No. Kafka 4.0 removed ZooKeeper completely and runs in KRaft mode, where brokers manage cluster metadata themselves. Older clusters you meet at work may still use ZooKeeper.

Polling with an updated_at filter misses deletes and collapses rapid changes into the last state. CDC reads the database transaction log, so every insert, update, and delete arrives in commit order without extra load on the tables.

Choose Kinesis when your stack is all AWS and you want almost no cluster operations. Choose Kafka or Amazon MSK when you need its larger ecosystem (Connect, Debezium, Schema Registry), long retention, or portability across clouds.