Glue Data Catalog
A central metadata repository that stores table definitions — column names, data types, and S3 locations — for your data lake. Athena, Redshift Spectrum, and EMR all use it to discover and query data without manual table definitions.
The AWS Glue Data Catalog is the central metadata repository for your data lake. It stores table definitions — column names, data types, partition information, and S3 file locations — so that Athena, Redshift Spectrum, and EMR can discover and query your data without manual table definitions.
The Problem It Solves
Without Data Catalog: New dataset lands in S3 Analyst must write: CREATE EXTERNAL TABLE orders (order_id STRING, amount DOUBLE, ...) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' LOCATION 's3://devops-raw/orders/' Every schema change → manual DDL update Every new analyst → writes the same CREATE TABLE again With Glue Data Catalog + Crawler: Crawler scans S3 → detects column names and types automatically Writes table definition to catalog Athena sees the table immediately — no CREATE TABLE needed Schema changes detected on next crawler run automaticallyCatalog Structure
Data Catalog└── Database: devops_analytics ├── Table: orders │ Columns: order_id (string), user_id (string), amount (double), city (string) │ Location: s3://devops-raw/orders/ │ Format: Parquet │ Partitioned by: year, month, day └── Table: users Columns: user_id (string), name (string), email (string) Location: s3://devops-raw/users/Services That Use the Catalog
Athena: runs SQL queries using catalog table definitions as the schemaRedshift Spectrum: extends Redshift to query S3 data lake tables via catalogAmazon EMR: Spark and Hive jobs discover datasets via catalogGlue ETL jobs: read catalog to understand source and target schemasAWS Lake Formation: governs row-level and column-level access to catalog tablesThe catalog is the connective tissue of your data lake — every analytics service reads from it.
Glue Crawlers
A Crawler scans your data sources, detects structure, and writes or updates table definitions:
Schedule: on demand, hourly, daily, or weeklySources: S3, RDS, DynamoDB, JDBC databases, MongoDBOn run: samples files → infers column names and types → creates or updates tableNew partition (new day of data): detected and added automatically on next crawlPartitioning — The Most Impactful Optimisation
Organise S3 files as: s3://bucket/orders/year=2024/month=01/day=15/file.parquet
The catalog stores partition locations. Athena uses them:
Query: WHERE year=2024 AND month=01 AND day=15Athena reads: only the matching partition (one day's data)vs scanning the entire dataset (years of data)Result: 100x less data scanned = 100x cheaper = much fasterGlue ETL Jobs
Beyond cataloging, Glue runs serverless Spark ETL jobs to transform data:
CSV files in S3 → Glue ETL job → Parquet files in S3Glue Crawler updates the catalog → Athena can query the Parquet immediatelyJob Bookmarks: track what was already processed — subsequent runs only process new data.
TipThe Glue Data Catalog is a one-time setup that pays dividends forever. Set up Crawlers from day one. Every new dataset that lands in S3 becomes queryable by all your analytics tools in minutes without any manual schema work.
Frequently Asked Questions
Why is the Glue Data Catalog described as central even though Athena, Redshift Spectrum, and EMR are separate services?
It stores table metadata (schema, partitions, S3 location) once, and any of those query engines can read the same catalog entry to know how to interpret raw files in S3 — Parquet, ORC, CSV, JSON — without each service needing its own separate schema registry. This means a single Glue Crawler run, or a manually defined table, makes that dataset immediately queryable from Athena, Redshift Spectrum, and EMR/Spark simultaneously.
What's a common mistake when relying on Glue Crawlers to keep the catalog up to date?
Running crawlers on a fixed schedule against frequently-changing or streaming data can miss newly added partitions between runs, causing queries in Athena to silently skip recent data rather than error out. For predictable partition schemes (like hourly date-based S3 prefixes), using `MSCK REPAIR TABLE` or the Glue Catalog API to register partitions programmatically as data lands is more reliable than waiting on a crawler's next scheduled pass.