Skip to main content

Glue Data Catalog

A central metadata repository that stores table definitions — column names, data types, and S3 locations — for your data lake. Athena, Redshift Spectrum, and EMR all use it to discover and query data without manual table definitions.

The AWS Glue Data Catalog is the central metadata repository for your data lake. It stores table definitions — column names, data types, partition information, and S3 file locations — so that Athena, Redshift Spectrum, and EMR can discover and query your data without manual table definitions.

The Problem It Solves

◈ DIAGRAM
Without Data Catalog:
New dataset lands in S3
Analyst must write:
CREATE EXTERNAL TABLE orders (order_id STRING, amount DOUBLE, ...)
ROW FORMAT DELIMITED FIELDS TERMINATED BY ','
LOCATION 's3://devops-raw/orders/'
Every schema change → manual DDL update
Every new analyst → writes the same CREATE TABLE again
With Glue Data Catalog + Crawler:
Crawler scans S3 → detects column names and types automatically
Writes table definition to catalog
Athena sees the table immediately — no CREATE TABLE needed
Schema changes detected on next crawler run automatically

Catalog Structure

◈ DIAGRAM
Data Catalog
└── Database: devops_analytics
├── Table: orders
│ Columns: order_id (string), user_id (string), amount (double), city (string)
│ Location: s3://devops-raw/orders/
│ Format: Parquet
│ Partitioned by: year, month, day
└── Table: users
Columns: user_id (string), name (string), email (string)
Location: s3://devops-raw/users/

Services That Use the Catalog

TEXT
Athena: runs SQL queries using catalog table definitions as the schema
Redshift Spectrum: extends Redshift to query S3 data lake tables via catalog
Amazon EMR: Spark and Hive jobs discover datasets via catalog
Glue ETL jobs: read catalog to understand source and target schemas
AWS Lake Formation: governs row-level and column-level access to catalog tables

The catalog is the connective tissue of your data lake — every analytics service reads from it.

Glue Crawlers

A Crawler scans your data sources, detects structure, and writes or updates table definitions:

◈ DIAGRAM
Schedule: on demand, hourly, daily, or weekly
Sources: S3, RDS, DynamoDB, JDBC databases, MongoDB
On run: samples files → infers column names and types → creates or updates table
New partition (new day of data): detected and added automatically on next crawl

Partitioning — The Most Impactful Optimisation

Organise S3 files as: s3://bucket/orders/year=2024/month=01/day=15/file.parquet

The catalog stores partition locations. Athena uses them:

TEXT
Query: WHERE year=2024 AND month=01 AND day=15
Athena reads: only the matching partition (one day's data)
vs scanning the entire dataset (years of data)
Result: 100x less data scanned = 100x cheaper = much faster

Glue ETL Jobs

Beyond cataloging, Glue runs serverless Spark ETL jobs to transform data:

◈ DIAGRAM
CSV files in S3 → Glue ETL job → Parquet files in S3
Glue Crawler updates the catalog → Athena can query the Parquet immediately

Job Bookmarks: track what was already processed — subsequent runs only process new data.

Tip

The Glue Data Catalog is a one-time setup that pays dividends forever. Set up Crawlers from day one. Every new dataset that lands in S3 becomes queryable by all your analytics tools in minutes without any manual schema work.

Frequently Asked Questions

Why is the Glue Data Catalog described as central even though Athena, Redshift Spectrum, and EMR are separate services?

It stores table metadata (schema, partitions, S3 location) once, and any of those query engines can read the same catalog entry to know how to interpret raw files in S3 — Parquet, ORC, CSV, JSON — without each service needing its own separate schema registry. This means a single Glue Crawler run, or a manually defined table, makes that dataset immediately queryable from Athena, Redshift Spectrum, and EMR/Spark simultaneously.

What's a common mistake when relying on Glue Crawlers to keep the catalog up to date?

Running crawlers on a fixed schedule against frequently-changing or streaming data can miss newly added partitions between runs, causing queries in Athena to silently skip recent data rather than error out. For predictable partition schemes (like hourly date-based S3 prefixes), using `MSCK REPAIR TABLE` or the Glue Catalog API to register partitions programmatically as data lands is more reliable than waiting on a crawler's next scheduled pass.