Skip to main content

Embeddings, Vector Databases, and AI Data Pipelines

Learn to turn documents into searchable meaning using embeddings, pgvector, and a real ingestion pipeline with metadata, dedup, and PII redaction.

~4 hours
9 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Keyword Search Fails acme-assist

At this point in the roadmap you have chosen a model and written the prompts for acme-assist, the help-centre chatbot you are building for acme-shop.

Generating Embeddings the Right Way

Choosing between a local model and a hosted API You can generate embeddings with an open-weight model running locally, or by calling a hosted...

Measuring How Close Two Meanings Are

Cosine similarity in plain terms Cosine similarity measures the angle between two vectors, not the distance between their tips.

Understanding What a Vector Database Stores

Why a Python loop stops scaling Comparing a query vector against 50 stored vectors in a loop is fine.

Choosing and Setting Up pgvector

Why pgvector is the default for this course Dedicated vector databases such as Pinecone, Weaviate, Qdrant, and Milvus exist for very large scale...

Designing the Ingestion Pipeline

Why "split and embed" is not a pipeline Many tutorials show two steps: split a document, embed the pieces.

Skills You'll Master

EMBEDDINGSVECTOR-DATABASEPGVECTORAI-DATA-PIPELINESEMANTIC-SEARCH

Curriculum Index9 topics

Career Impact

Roles that use the skills in this module.

  • AI Engineer

  • MLOps Engineer

  • Platform Engineer

  • Data Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Keyword search matches the words in a query against the words in a document. Embedding search compares the meaning of the two, so a query like "my payout is stuck" can find a page titled "Resolving Delayed Settlement Issues" even though they share no words.

Not at first. If you already run Postgres, the pgvector extension handles many workloads. Move to a dedicated vector database when measured latency, scale, or operational needs justify the extra system.

Vectors from different models cannot be compared. Recording the model per vector lets you find exactly which rows need re-embedding when you switch models, instead of silently mixing incompatible vectors.

No. Redaction removes known patterns from the text you store. It does not stop one user's query from retrieving another user's document, so you also need access control on retrieval.