Skip to main content

Cloud Platforms for Data Engineers

Learn the AWS services data engineers use daily: an S3 data lake, the Glue catalog and jobs, Athena, Redshift, and least-privilege IAM for pipelines.

~4 hours
13 Topics
Hands-on Scenarios

What You'll Learn

Understanding What This Module Covers and How the Pieces Fit

It is month three on acme-shop's data team.

Designing S3 as a Data Lake, Not a File Dump

A bucket with a thousand CSV files in one flat folder is not a data lake. It is a junk drawer with a price tag.

Moving Data Safely with S3 Events and a Quarantine Zone

Bad files do not announce themselves. A source system renames a column on a Tuesday, and by Wednesday the dashboard is wrong with no error anywhere.

Cataloguing Data with the Glue Data Catalog and Crawlers

Files in S3 are just bytes until something records what they mean. The catalog is that record.

Transforming Data with Glue ETL Jobs

Raw CSV is almost never what analysts should query.

Querying Data with Athena

Athena is where most analysts meet your data lake.

Skills You'll Master

AWSS3-DATA-LAKEAWS-GLUEATHENADATA-ENGINEERING

Curriculum Index13 topics

1

Understanding What This Module Covers and How the Pieces Fit

It is month three on acme-shop's data team.

2

Designing S3 as a Data Lake, Not a File Dump

A bucket with a thousand CSV files in one flat folder is not a data lake. It is a junk drawer with a price tag.

3

Moving Data Safely with S3 Events and a Quarantine Zone

Bad files do not announce themselves. A source system renames a column on a Tuesday, and by Wednesday the dashboard is...

4

Cataloguing Data with the Glue Data Catalog and Crawlers

Files in S3 are just bytes until something records what they mean. The catalog is that record.

5

Transforming Data with Glue ETL Jobs

Raw CSV is almost never what analysts should query.

6

Querying Data with Athena

Athena is where most analysts meet your data lake.

7

Choosing Redshift for Repeated Heavy Queries

Athena is perfect until the same expensive dashboard query runs 400 times a day.

8

Streaming at a Glance with Kinesis and Amazon Data Firehose

Not every data source arrives as a daily file.

9

Securing Data Pipelines with IAM Roles

Every service in this module talks to another. The question is how they prove who they are.

10

Provisioning the Data Lake Baseline with Terraform

A data lake you built by clicking through the console cannot be rebuilt, reviewed, or torn down reliably.

11

Troubleshooting a Query That Scans Everything

A teammate reports that a query for one day of acme-shop orders costs far more than expected, even though the table is...

12

Hands-On Lab: Building and Querying a Cloud Data Lake

You will build acme-shop's data lake with Terraform, run a Glue job on the orders data, query it with Athena, break the...

13

Quick Reference and Common Mistakes

Quick Reference Common Mistakes Storing data in S3 with no partitioning forces Athena to scan the whole dataset on...

Career Impact

Roles that use the skills in this module.

  • Data Engineer

  • Platform Engineer

  • Cloud Engineer

  • DevOps Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

No, but you should be comfortable with an AWS account, the CLI, and what an IAM role is. The Cloud Engineer roadmap steps Cloud Computing and AWS Fundamentals and AWS Core Services teach that foundation in depth. This module only covers the data-specific slice.

S3 stores the data files. The Glue Data Catalog stores metadata about them: column names, types, partitions, and the S3 location. Athena is the query engine that reads the catalog to know what to look for, then reads the matching bytes from S3.

Parquet stores each column together and compresses it, so a query that needs three columns out of thirty reads only those three. CSV forces the engine to read every byte of every row. On the acme-shop orders file used in this module, Parquet is already less than half the size of the CSV.

Choose Athena for occasional or unpredictable queries, because you pay only per query. Choose Redshift when the same heavy queries run often for many users, where a steady compute cost can beat paying per byte scanned every time. Measure your own workload before deciding.