Cloud Platforms for Data Engineers
Learn the AWS services data engineers use daily: an S3 data lake, the Glue catalog and jobs, Athena, Redshift, and least-privilege IAM for pipelines.
What You'll Learn
Understanding What This Module Covers and How the Pieces Fit
It is month three on acme-shop's data team.
Designing S3 as a Data Lake, Not a File Dump
A bucket with a thousand CSV files in one flat folder is not a data lake. It is a junk drawer with a price tag.
Moving Data Safely with S3 Events and a Quarantine Zone
Bad files do not announce themselves. A source system renames a column on a Tuesday, and by Wednesday the dashboard is wrong with no error anywhere.
Cataloguing Data with the Glue Data Catalog and Crawlers
Files in S3 are just bytes until something records what they mean. The catalog is that record.
Transforming Data with Glue ETL Jobs
Raw CSV is almost never what analysts should query.
Querying Data with Athena
Athena is where most analysts meet your data lake.
Skills You'll Master
Curriculum Index13 topics
Understanding What This Module Covers and How the Pieces Fit
It is month three on acme-shop's data team.
Designing S3 as a Data Lake, Not a File Dump
A bucket with a thousand CSV files in one flat folder is not a data lake. It is a junk drawer with a price tag.
Moving Data Safely with S3 Events and a Quarantine Zone
Bad files do not announce themselves. A source system renames a column on a Tuesday, and by Wednesday the dashboard is...
Cataloguing Data with the Glue Data Catalog and Crawlers
Files in S3 are just bytes until something records what they mean. The catalog is that record.
Transforming Data with Glue ETL Jobs
Raw CSV is almost never what analysts should query.
Querying Data with Athena
Athena is where most analysts meet your data lake.
Choosing Redshift for Repeated Heavy Queries
Athena is perfect until the same expensive dashboard query runs 400 times a day.
Streaming at a Glance with Kinesis and Amazon Data Firehose
Not every data source arrives as a daily file.
Securing Data Pipelines with IAM Roles
Every service in this module talks to another. The question is how they prove who they are.
Provisioning the Data Lake Baseline with Terraform
A data lake you built by clicking through the console cannot be rebuilt, reviewed, or torn down reliably.
Troubleshooting a Query That Scans Everything
A teammate reports that a query for one day of acme-shop orders costs far more than expected, even though the table is...
Hands-On Lab: Building and Querying a Cloud Data Lake
You will build acme-shop's data lake with Terraform, run a Glue job on the orders data, query it with Athena, break the...
Quick Reference and Common Mistakes
Quick Reference Common Mistakes Storing data in S3 with no partitioning forces Athena to scan the whole dataset on...
Career Impact
Roles that use the skills in this module.
Data Engineer
Platform Engineer
Cloud Engineer
DevOps Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
No, but you should be comfortable with an AWS account, the CLI, and what an IAM role is. The Cloud Engineer roadmap steps Cloud Computing and AWS Fundamentals and AWS Core Services teach that foundation in depth. This module only covers the data-specific slice.
S3 stores the data files. The Glue Data Catalog stores metadata about them: column names, types, partitions, and the S3 location. Athena is the query engine that reads the catalog to know what to look for, then reads the matching bytes from S3.
Parquet stores each column together and compresses it, so a query that needs three columns out of thirty reads only those three. CSV forces the engine to read every byte of every row. On the acme-shop orders file used in this module, Parquet is already less than half the size of the CSV.
Choose Athena for occasional or unpredictable queries, because you pay only per query. Choose Redshift when the same heavy queries run often for many users, where a steady compute cost can beat paying per byte scanned every time. Measure your own workload before deciding.