Skip to main content

RAG in Production: Build, Debug, Improve

Learn to build a production RAG pipeline: chunking, hybrid search, reranking, parent-child retrieval, citations, and how to debug wrong answers.

~4 hours
13 Topics
Hands-on Scenarios

What You'll Learn

Understanding What RAG Solves and When to Use It

A customer in Jaipur asks the Acme Shop chatbot how long she has to return a laptop. The bot answers with confidence: 30 days.

Preparing Documents and Metadata for Retrieval

Retrieval can only find what you indexed, so the quality of your cleaned text and metadata sets a ceiling that no clever search technique can lift.

Choosing a Chunking Strategy

A chunk is the unit of text you embed and retrieve, and choosing its size and boundaries decides whether the right fact is findable and whether it...

Running Vector Search with Metadata Filters

Vector search is the first retrieval stage, and it only works well when queries and documents are embedded the same way and when filters remove...

Adding Hybrid Search with Keyword Matching

Vector search understands meaning but is weak on exact strings, so adding keyword search catches the product codes, form numbers, and names that...

Improving Retrieval with Query Rewriting and Reranking

Even good first-stage search returns a rough ordering, so rewriting the query before search and reranking the results after it are the two cheapest...

Skills You'll Master

RAGRETRIEVALRERANKINGHYBRID-SEARCHPGVECTOR

Curriculum Index13 topics

1

Understanding What RAG Solves and When to Use It

A customer in Jaipur asks the Acme Shop chatbot how long she has to return a laptop.

2

Preparing Documents and Metadata for Retrieval

Retrieval can only find what you indexed, so the quality of your cleaned text and metadata sets a ceiling that no...

3

Choosing a Chunking Strategy

A chunk is the unit of text you embed and retrieve, and choosing its size and boundaries decides whether the right fact...

4

Running Vector Search with Metadata Filters

Vector search is the first retrieval stage, and it only works well when queries and documents are embedded the same way...

5

Adding Hybrid Search with Keyword Matching

Vector search understands meaning but is weak on exact strings, so adding keyword search catches the product codes...

6

Improving Retrieval with Query Rewriting and Reranking

Even good first-stage search returns a rough ordering, so rewriting the query before search and reranking the results...

7

Using Parent-Child Retrieval for Fuller Context

Small chunks are best for finding a fact and large chunks are best for answering a question, and parent-child retrieval...

8

Generating Grounded Answers with Citations

The generation step is where a good retrieval result can still be ruined, so the prompt has to force the model to...

9

Debugging Wrong Answers in a RAG System

When an answer is wrong, the fastest fix comes from finding which stage lost the right passage, and you can do that by...

10

Measuring Retrieval Quality with a Golden Set

Without numbers, every change feels like an improvement, so a small fixed set of questions is the cheapest way to know...

11

Building the Acme Policy Chatbot End to End

This lab builds acme-assist, a policy chatbot that starts with a plain vector search and improves step by step, with a...

12

Using the RAG Quick Reference

Keep this section open when you build or debug a pipeline, because it condenses the decisions from every topic into one...

13

Avoiding Common RAG Mistakes

These seven mistakes cause most failed RAG projects, and every one of them is cheaper to prevent than to debug.

Career Impact

Roles that use the skills in this module.

  • AI Engineer

  • MLOps Engineer

  • Platform Engineer

  • DevOps Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

RAG stands for Retrieval Augmented Generation. Your code first searches your own documents for passages that relate to the question, then gives only those passages to a language model and asks it to answer from them. It works like an open-book exam: the model does not need to remember the facts because it can read them.

Yes, for most real systems. Sending hundreds of pages with every question is slow and expensive, and every user would see every document. Retrieval also lets you filter by permission and cite the exact source. Pasting the whole text is fine for one short document.

There is no universal size. Start by splitting along the structure of the document, such as sections and paragraphs, and keep each chunk to a few hundred tokens. Then measure on real questions and adjust, because the right size depends on how your documents are written.

Most of the time retrieval failed, not the model. The correct passage was never found, was ranked too low, or was filtered out. Check each stage in order: is the fact in the index, in the candidate list, in the final top results, and in the prompt?

Use RAG when answers must come from documents that change or must be cited. Use fine-tuning to change behaviour, such as tone or output format. Many teams try better prompting first, then RAG, and fine-tune only when the first two are not enough.