Skip to main content

Apache Spark for Batch Processing

Learn PySpark for data too big for one machine: DataFrames, partitions, shuffles, joins, the Spark UI, and the fixes that make slow jobs fast.

~4 hours
13 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Spark Exists

It is the first morning after acme-shop's biggest sale.

Understanding Spark Architecture

Every Spark application is a small team: one coordinator and several workers.

Working with PySpark DataFrames

PySpark is Spark's Python API. A DataFrame looks like a pandas DataFrame, but its rows are split across partitions and processed by several executors...

Understanding Partitions, Shuffles, and Lazy Evaluation

Three ideas explain most Spark performance problems. Spark is not "pandas but distributed": once you see these three, slow jobs stop being mysterious.

Reading Plans and the Spark UI

When a job is slow, do not guess. Spark shows you exactly what it planned and what actually happened, in two places.

Fixing Data Skew and Using Adaptive Query Execution

What skew is and how to spot it Parallel work is only fast when it is balanced. Skew means one partition holds far more rows than the others.

Skills You'll Master

SPARKPYSPARKDATABRICKSBATCH-PROCESSINGBIG-DATA

Curriculum Index13 topics

1

Understanding Why Spark Exists

It is the first morning after acme-shop's biggest sale.

2

Understanding Spark Architecture

Every Spark application is a small team: one coordinator and several workers.

3

Working with PySpark DataFrames

PySpark is Spark's Python API. A DataFrame looks like a pandas DataFrame, but its rows are split across partitions and...

4

Understanding Partitions, Shuffles, and Lazy Evaluation

Three ideas explain most Spark performance problems.

5

Reading Plans and the Spark UI

When a job is slow, do not guess. Spark shows you exactly what it planned and what actually happened, in two places.

6

Fixing Data Skew and Using Adaptive Query Execution

What skew is and how to spot it Parallel work is only fast when it is balanced.

7

Optimising Spark Jobs

Broadcast joins: skip the shuffle for small tables If one side of a join is small enough to fit in memory on every...

8

Running Spark with Databricks and Managed Services

Running a Spark cluster yourself means patching, sizing, and monitoring machines.

9

Troubleshooting a Join That Never Finishes

A teammate hands you this job. "It worked on a sample, but it has run for two hours on the full data." Diagnose before...

10

Running the Hands-on Lab

You will generate 20 million acme-shop orders on your own laptop, run the same join two ways, read the plan and the...

11

Extending the Lab: Running the Job on AWS Glue

This extension is optional. It shows the same kind of job running on managed serverless Spark.

12

Understanding What You Built and What Comes Next

What you built You generated 20 million acme-shop orders on your laptop, ran the same join as a shuffle join and a...

13

Reviewing the Quick Reference and Common Mistakes

Quick reference Common mistakes Relying on schema inference in a recurring pipeline makes Spark scan the file on every...

Career Impact

Roles that use the skills in this module.

  • Data Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Use Spark when the work no longer fits on one machine or no longer finishes inside its time window. If pandas, DuckDB, or your warehouse handles it comfortably, Spark only adds cluster cost and complexity.

A shuffle moves rows between machines so that rows with the same key end up together. It writes data to disk, sends it over the network, and reads it back, and it happens for joins, groupBy, and sorts.

Transformations like filter and join only build a plan. Nothing runs until an action such as write, count, or show asks for a result, which lets Spark optimise the whole plan at once.

No. Spark runs in local mode on a laptop, using your CPU cores as executors. That is enough to read plans, use the Spark UI, and see shuffles, broadcast joins, and skew.