Skip to main content

Data Governance and Catalog

Learn practical data governance: classify data, catalog and lineage, role and row access, masking, retention, and erasure under GDPR and DPDP.

~3 hours
13 Topics
Hands-on Scenarios

What You'll Learn

An Advanced Competency, Not a Gate You Must Pass First

You have joined acme-shop's data team and built pipelines, a star schema, tested dbt models, and quality gates.

Data Classification - Know What You Are Protecting Before You Protect It

Data classification is the practice of labeling every dataset by sensitivity, so governance decisions are made consistently instead of case by case.

Data Catalog - The Searchable Inventory of What You Have

A data catalog is a searchable inventory of every data asset: each table, its owner, a description, and where to find it.

Data Lineage - Tracing Where Data Came From

Data lineage is the traceable path data takes from its source, through every transformation, to where it lands.

Role-Based Access Control - Who Can See What

Role-based access control (RBAC) grants access by role, such as analyst or finance, instead of person by person.

Data Masking, Pseudonymisation, and Anonymisation - Protecting Data in Dev and Test Environments

Production data often gets copied into dev or staging so engineers can test on realistic data, and that is how a real customer email ends up on a...

Skills You'll Master

DATA-GOVERNANCEDATA-CATALOGDATA-LINEAGEDATA-PRIVACYACCESS-CONTROL

Curriculum Index13 topics

1

An Advanced Competency, Not a Gate You Must Pass First

You have joined acme-shop's data team and built pipelines, a star schema, tested dbt models, and quality gates.

2

Data Classification - Know What You Are Protecting Before You Protect It

Data classification is the practice of labeling every dataset by sensitivity, so governance decisions are made...

3

Data Catalog - The Searchable Inventory of What You Have

A data catalog is a searchable inventory of every data asset: each table, its owner, a description, and where to find...

4

Data Lineage - Tracing Where Data Came From

Data lineage is the traceable path data takes from its source, through every transformation, to where it lands.

5

Role-Based Access Control - Who Can See What

Role-based access control (RBAC) grants access by role, such as analyst or finance, instead of person by person.

6

Data Masking, Pseudonymisation, and Anonymisation - Protecting Data in Dev and Test Environments

Production data often gets copied into dev or staging so engineers can test on realistic data, and that is how a real...

7

GDPR Essentials for Data Engineers - Practical Implementation, Not a Legal Deep Dive

GDPR is the European Union's data protection law, and its patterns have become the default expectation for any company...

8

Understanding India's DPDP Act for Data Engineers

The Digital Personal Data Protection Act, 2023 (DPDP Act) is India's data protection law.

9

Troubleshooting Scenario - Find and Fix the Governance Gap

acme-shop's erasure script deletes a customer from raw and the marts whenever a request comes in.

10

The Governance Loop - One Diagram to Remember

Every technique in this module is a step in one repeating cycle that keeps new data trustworthy for as long as it...

11

Hands-On Lab

📌 Remember: this lab is free and runs entirely on your laptop. It needs Docker, and about 15 minutes of machine time.

12

Quick Reference

Table: Task, Tool or pattern

13

Common Mistakes

Treating the data catalog as a one-time documentation project means it is accurate on launch day and quietly wrong a...

Career Impact

Roles that use the skills in this module.

  • Data Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Not on day one. Reliable pipelines, good modeling, and data quality come first. Governance is the layer you add once those exist, and it becomes part of the job as soon as your pipelines touch personal data.

Masking hides a value from a viewer, often reversibly. Pseudonymisation swaps an identifier for a token, but the person can still be re-identified with the right mapping. Anonymisation means the person can no longer reasonably be identified at all.

Not always. Some records, such as completed financial transactions, may have to be kept for legal reasons. In that case the usual approach is to remove or anonymise the personal fields and keep the minimum record the law requires. Legal and privacy teams decide which fields.

The Digital Personal Data Protection Act, 2023 applies to processing of digital personal data in India, and in some cases to processing outside India that offers goods or services to people in India. Your legal team confirms how it applies. As a data engineer you build the systems that meet it.