Skip to main content

Real-Time Stream Processing with Apache Flink

Learn Flink SQL to turn Kafka streams into live results: windows, event time, watermarks, state, and checkpoints, at a working level.

~3 hours
14 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Stream Processing Engines Exist

It is 2:14 AM, and a customer pays for an acme-shop order with UPI.

Choosing Flink SQL Over the DataStream API

Flink offers three ways to write jobs: the DataStream API (Java, Scala, or Python, full control), the Table API, and Flink SQL.

Reading from Kafka and Writing a Windowed Aggregation

Every Flink SQL streaming job has the same three steps: declare a source table, declare a sink table, then run a query that reads one and writes the...

Understanding Windowing at a Conceptual Level

A stream never ends, so "give me the total" has no meaning on its own: total of what, up to when?

Understanding Event Time, Processing Time, and Watermarks

Every event has two timestamps. Event time is when it actually happened. Processing time is when Flink happens to see it.

Debugging a Window That Stops Emitting Results

This scenario happens in production far more often than a missing watermark clause. The Symptom A job with a correct watermark ran fine for weeks.

Skills You'll Master

FLINKFLINK-SQLSTREAM-PROCESSINGWATERMARKSKAFKA

Curriculum Index14 topics

1

Understanding Why Stream Processing Engines Exist

It is 2:14 AM, and a customer pays for an acme-shop order with UPI.

2

Choosing Flink SQL Over the DataStream API

Flink offers three ways to write jobs: the DataStream API (Java, Scala, or Python, full control), the Table API, and...

3

Reading from Kafka and Writing a Windowed Aggregation

Every Flink SQL streaming job has the same three steps: declare a source table, declare a sink table, then run a query...

4

Understanding Windowing at a Conceptual Level

A stream never ends, so "give me the total" has no meaning on its own: total of what, up to when?

5

Understanding Event Time, Processing Time, and Watermarks

Every event has two timestamps. Event time is when it actually happened.

6

Debugging a Window That Stops Emitting Results

This scenario happens in production far more often than a missing watermark clause.

7

Understanding State

Every windowed total you have written implies Flink is remembering something between events.

8

Understanding Checkpointing and Processing Guarantees

Your job has run for six hours, holding partial totals. The machine crashes.

9

Matching Kafka Partitions to Flink Parallelism

Parallelism is how many parallel subtasks Flink runs for part of a job.

10

Recognising Streaming Joins and Streaming into a Lake

These are awareness topics. You should know the words, and not write them from scratch yet.

11

Choosing Where Flink Fits

You now have three tools that look alike but solve different problems.

12

Running the Hands-On Lab

You will run Kafka and Flink locally with Docker, stream acme-shop order events, and watch a 30-second window close...

13

Quick Reference and Common Mistakes

Quick Reference Common Mistakes Starting with the DataStream API.

14

What You Built and What Comes Next

You built a live revenue-per-category job for acme-shop with Flink SQL, saw a window close only when the watermark...

Career Impact

Roles that use the skills in this module.

  • Data Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Kafka stores and delivers streams of events reliably. Flink reads those events and computes something over them, such as totals per window or alerts. Kafka moves the data, and Flink does the thinking.

Start with Flink SQL. Most streaming ETL jobs, such as filters, windowed aggregations, and simple joins, can be written in SQL, which is easier to read and maintain. Use the DataStream API only for custom logic that SQL cannot express.

A window emits only when the watermark passes its end time, and the watermark advances only when Flink sees newer event times. A missing watermark, no newer events, or an idle Kafka partition are the usual causes.

Only if the business needs results within seconds. Spark is simpler for bounded batch jobs, and many teams run both. Choose Flink when freshness is a real requirement, not because streaming sounds more impressive.