Apache Spark for Batch Processing
Learn PySpark for data too big for one machine: DataFrames, partitions, shuffles, joins, the Spark UI, and the fixes that make slow jobs fast.
What You'll Learn
Understanding Why Spark Exists
It is the first morning after acme-shop's biggest sale.
Understanding Spark Architecture
Every Spark application is a small team: one coordinator and several workers.
Working with PySpark DataFrames
PySpark is Spark's Python API. A DataFrame looks like a pandas DataFrame, but its rows are split across partitions and processed by several executors...
Understanding Partitions, Shuffles, and Lazy Evaluation
Three ideas explain most Spark performance problems. Spark is not "pandas but distributed": once you see these three, slow jobs stop being mysterious.
Reading Plans and the Spark UI
When a job is slow, do not guess. Spark shows you exactly what it planned and what actually happened, in two places.
Fixing Data Skew and Using Adaptive Query Execution
What skew is and how to spot it Parallel work is only fast when it is balanced. Skew means one partition holds far more rows than the others.
Skills You'll Master
Curriculum Index13 topics
Understanding Why Spark Exists
It is the first morning after acme-shop's biggest sale.
Understanding Spark Architecture
Every Spark application is a small team: one coordinator and several workers.
Working with PySpark DataFrames
PySpark is Spark's Python API. A DataFrame looks like a pandas DataFrame, but its rows are split across partitions and...
Understanding Partitions, Shuffles, and Lazy Evaluation
Three ideas explain most Spark performance problems.
Reading Plans and the Spark UI
When a job is slow, do not guess. Spark shows you exactly what it planned and what actually happened, in two places.
Fixing Data Skew and Using Adaptive Query Execution
What skew is and how to spot it Parallel work is only fast when it is balanced.
Optimising Spark Jobs
Broadcast joins: skip the shuffle for small tables If one side of a join is small enough to fit in memory on every...
Running Spark with Databricks and Managed Services
Running a Spark cluster yourself means patching, sizing, and monitoring machines.
Troubleshooting a Join That Never Finishes
A teammate hands you this job. "It worked on a sample, but it has run for two hours on the full data." Diagnose before...
Running the Hands-on Lab
You will generate 20 million acme-shop orders on your own laptop, run the same join two ways, read the plan and the...
Extending the Lab: Running the Job on AWS Glue
This extension is optional. It shows the same kind of job running on managed serverless Spark.
Understanding What You Built and What Comes Next
What you built You generated 20 million acme-shop orders on your laptop, ran the same join as a shuffle join and a...
Reviewing the Quick Reference and Common Mistakes
Quick reference Common mistakes Relying on schema inference in a recurring pipeline makes Spark scan the file on every...
Career Impact
Roles that use the skills in this module.
Data Engineer
Platform Engineer
Next Modules
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
Use Spark when the work no longer fits on one machine or no longer finishes inside its time window. If pandas, DuckDB, or your warehouse handles it comfortably, Spark only adds cluster cost and complexity.
A shuffle moves rows between machines so that rows with the same key end up together. It writes data to disk, sends it over the network, and reads it back, and it happens for joins, groupBy, and sorts.
Transformations like filter and join only build a plan. Nothing runs until an action such as write, count, or show asks for a result, which lets Spark optimise the whole plan at once.
No. Spark runs in local mode on a laptop, using your CPU cores as executors. That is enough to read plans, use the Spark UI, and see shuffles, broadcast joins, and skew.