RAG for Ops Knowledge
Learn to ground ops agents in your own runbooks and postmortems: chunk, embed, retrieve with citations, and know when live data beats retrieval.
What You'll Learn
Understanding Why an Ops Agent Needs Your Knowledge
It is 2 AM at acme-shop. Checkout is returning 502 and the on-call engineer asks the AI assistant what to do.
Understanding Embeddings and Similarity
An embedding is a list of numbers that represents the meaning of a piece of text, produced by an embedding model.
Choosing a Vector Database
A vector database stores vectors alongside the original text and metadata, and answers "which stored vectors are closest to this one" quickly.
Chunking Runbooks and Postmortems
Chunking means splitting each document into smaller pieces, and each piece gets its own vector.
Retrieving with Hybrid Search and Reciprocal Rank Fusion
Retrieval is where most RAG systems succeed or fail.
Answering with Citations and Saying "No Runbook Found"
The last step turns retrieved chunks into an answer an engineer can trust.
Skills You'll Master
Curriculum Index10 topics
Understanding Why an Ops Agent Needs Your Knowledge
It is 2 AM at acme-shop. Checkout is returning 502 and the on-call engineer asks the AI assistant what to do.
Understanding Embeddings and Similarity
An embedding is a list of numbers that represents the meaning of a piece of text, produced by an embedding model.
Choosing a Vector Database
A vector database stores vectors alongside the original text and metadata, and answers "which stored vectors are...
Chunking Runbooks and Postmortems
Chunking means splitting each document into smaller pieces, and each piece gets its own vector.
Retrieving with Hybrid Search and Reciprocal Rank Fusion
Retrieval is where most RAG systems succeed or fail.
Answering with Citations and Saying "No Runbook Found"
The last step turns retrieved chunks into an answer an engineer can trust.
Knowing When Not to Use RAG
RAG is for knowledge that changes slowly: procedures, history, ownership, architecture.
Debugging and Evaluating a RAG System
A demo that works on three questions says nothing about a 3 AM incident.
Hands-On Lab: A Runbook Assistant for acme-shop
📌 Remember: this lab is free. Retrieval needs only Python and about 2 GB of disk for the local embedding model.
Quick Reference and Common Mistakes
Quick reference Common mistakes Indexing documents nobody has reviewed means the assistant quotes whatever the runbook...
Career Impact
Roles that use the skills in this module.
- High Demand
AIOps Engineer
₹18L - ₹35L a year
- High Demand
AI Platform Engineer
₹20L - ₹40L a year
- Growing Fast
DevOps Engineer (AI)
₹15L - ₹28L a year
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
RAG stands for retrieval-augmented generation. The system first searches your own documents, then gives the best matches to the model so it answers from your runbooks instead of from general knowledge. Without it, an agent cannot know how your team handles your systems.
Runbooks change often, and retraining is slow, costly, and hard to undo. With RAG you re-index a changed file and the next question already sees it. Fine-tuning is better for changing how a model behaves, not for teaching it facts that change.
No. RAG searches documents that were indexed earlier, so it is for slow-changing knowledge such as procedures and past incidents. Live questions need live tools like Prometheus queries or the Kubernetes API.
Tell the model to answer only from the retrieved excerpts, return a fixed 'no runbook found' message when they do not answer the question, and check relevance before calling the model at all. Showing citations lets engineers verify every answer.