Skip to main content

CI/CD and Testing for Data Pipelines

Learn to test data pipelines with pytest and dbt, block bad pull requests in CI, deploy Airflow DAGs with GitHub Actions, and roll back safely.

~3 hours
8 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Data Pipelines Need CI/CD

The Monday morning revenue dashboard Most data incidents are not outages.

Testing Transformation Logic with pytest

What to unit test in a data pipeline Unit tests check small pieces of logic without touching a real database.

Running dbt Build and Tests in CI

What dbt build does in a pull request dbt build runs your seeds, models, snapshots and tests together, in dependency order.

Enforcing Data Contracts in Pull Requests

What a data contract protects A data contract is a promise about what a dataset looks like: its columns, their types, and rules about the values.

Deploying Airflow DAGs with GitHub Actions

Test that the DAG loads before it ships Airflow reads every Python file in the DAG folder.

Rolling Back a Pipeline That Produced Bad Data

Two problems, two fixes When bad data ships there are two separate problems: the bad logic is still running, and the bad rows are already in your...

Skills You'll Master

DATA-ENGINEERINGCI-CDDBTAIRFLOWGITHUB-ACTIONSPYTEST

Curriculum Index8 topics

Career Impact

Roles that use the skills in this module.

  • Data Engineer

  • Platform Engineer

  • DevOps Engineer

  • Release Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

SQL changes can silently change the numbers that finance, product and leadership rely on. CI runs the change against real tests before it merges, so a broken join or a renamed column is caught in a pull request and not in a Monday morning dashboard. CD then makes releases repeatable instead of copying files by hand.

Slim CI builds and tests only the models you changed and the models that depend on them, instead of the whole project. It compares your branch to the manifest from the last production run and defers unchanged models to production tables. This keeps pull request checks fast on large projects.

A data contract is a written, enforced agreement about the shape of a dataset: column names, data types and rules. In dbt, a model contract makes the build fail if the model's output does not match what you promised downstream users. It turns a surprise breaking change into a failed pull request.

At minimum, load the DAG folder in a test and assert there are no import errors, since a syntax or import error can stop the whole folder from loading. Then unit test the Python functions your tasks call, and run the transformations against a small test dataset. Deploy to a dev environment first and promote to production only after it runs cleanly.

Roll back in two parts. First revert the code and redeploy, so the bad logic stops running. Then repair the data by rebuilding the affected models, restoring from a snapshot or time travel if your platform has one, and re-running the affected dates. Building into a new table and swapping it in after checks makes the repair much safer.