CI/CD and Testing for Data Pipelines
Learn to test data pipelines with pytest and dbt, block bad pull requests in CI, deploy Airflow DAGs with GitHub Actions, and roll back safely.
What You'll Learn
Understanding Why Data Pipelines Need CI/CD
The Monday morning revenue dashboard Most data incidents are not outages.
Testing Transformation Logic with pytest
What to unit test in a data pipeline Unit tests check small pieces of logic without touching a real database.
Running dbt Build and Tests in CI
What dbt build does in a pull request dbt build runs your seeds, models, snapshots and tests together, in dependency order.
Enforcing Data Contracts in Pull Requests
What a data contract protects A data contract is a promise about what a dataset looks like: its columns, their types, and rules about the values.
Deploying Airflow DAGs with GitHub Actions
Test that the DAG loads before it ships Airflow reads every Python file in the DAG folder.
Rolling Back a Pipeline That Produced Bad Data
Two problems, two fixes When bad data ships there are two separate problems: the bad logic is still running, and the bad rows are already in your...
Skills You'll Master
Curriculum Index8 topics
Understanding Why Data Pipelines Need CI/CD
The Monday morning revenue dashboard Most data incidents are not outages.
Testing Transformation Logic with pytest
What to unit test in a data pipeline Unit tests check small pieces of logic without touching a real database.
Running dbt Build and Tests in CI
What dbt build does in a pull request dbt build runs your seeds, models, snapshots and tests together, in dependency...
Enforcing Data Contracts in Pull Requests
What a data contract protects A data contract is a promise about what a dataset looks like: its columns, their types...
Deploying Airflow DAGs with GitHub Actions
Test that the DAG loads before it ships Airflow reads every Python file in the DAG folder.
Rolling Back a Pipeline That Produced Bad Data
Two problems, two fixes When bad data ships there are two separate problems: the bad logic is still running, and the...
Running the Lab: A Pull Request That CI Blocks
Lab overview You will build a small orders pipeline with a pytest suite and a dbt project, wire it to GitHub Actions...
Quick Reference and Common Mistakes
Quick reference Common mistakes Running CI against production is the most dangerous one.
Career Impact
Roles that use the skills in this module.
Data Engineer
Platform Engineer
DevOps Engineer
Release Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
SQL changes can silently change the numbers that finance, product and leadership rely on. CI runs the change against real tests before it merges, so a broken join or a renamed column is caught in a pull request and not in a Monday morning dashboard. CD then makes releases repeatable instead of copying files by hand.
Slim CI builds and tests only the models you changed and the models that depend on them, instead of the whole project. It compares your branch to the manifest from the last production run and defers unchanged models to production tables. This keeps pull request checks fast on large projects.
A data contract is a written, enforced agreement about the shape of a dataset: column names, data types and rules. In dbt, a model contract makes the build fail if the model's output does not match what you promised downstream users. It turns a surprise breaking change into a failed pull request.
At minimum, load the DAG folder in a test and assert there are no import errors, since a syntax or import error can stop the whole folder from loading. Then unit test the Python functions your tasks call, and run the transformations against a small test dataset. Deploy to a dev environment first and promote to production only after it runs cleanly.
Roll back in two parts. First revert the code and redeploy, so the bad logic stops running. Then repair the data by rebuilding the affected models, restoring from a snapshot or time travel if your platform has one, and re-running the affected dates. Building into a new table and swapping it in after checks makes the repair much safer.