AI Evaluation and Testing
Learn to prove your AI system works: golden datasets, LLM-as-judge, retrieval and faithfulness metrics, agent evaluation, and release gates in CI.
What You'll Learn
Understanding Why AI Systems Need Their Own Testing
The acme-assist demo went well. Twenty questions, twenty fluent answers, and the support lead at acme-shop nodded all the way through.
Building a Golden Dataset
Before any metric, you need something to measure against. A fixed, trusted set of cases is the foundation for everything else in this module.
Scoring Answers with Rules and LLM Judges
You now have cases. Next you need to decide whether an answer passes, and the cheapest reliable check is usually not a model at all.
Measuring Retrieval Quality in RAG
A RAG system can fail in two independent ways, and a wrong final answer does not say which one happened.
Measuring Generation Quality in RAG
Once retrieval looks healthy, the remaining question is what the model did with the chunks it was given.
Evaluating AI Agents by Trajectory and Cost
An agent takes a sequence of actions, so the final result can look right even when the path was wasteful, risky, or lucky.
Skills You'll Master
Curriculum Index10 topics
Understanding Why AI Systems Need Their Own Testing
The acme-assist demo went well. Twenty questions, twenty fluent answers, and the support lead at acme-shop nodded all...
Building a Golden Dataset
Before any metric, you need something to measure against.
Scoring Answers with Rules and LLM Judges
You now have cases. Next you need to decide whether an answer passes, and the cheapest reliable check is usually not a...
Measuring Retrieval Quality in RAG
A RAG system can fail in two independent ways, and a wrong final answer does not say which one happened.
Measuring Generation Quality in RAG
Once retrieval looks healthy, the remaining question is what the model did with the chunks it was given.
Evaluating AI Agents by Trajectory and Cost
An agent takes a sequence of actions, so the final result can look right even when the path was wasteful, risky, or...
Building Regression Tests and Release Gates
Evaluation before launch finds the problems you expected.
Monitoring Drift and Avoiding Evaluation Leakage
Offline evaluation covers what you thought to test.
Running the Hands-On Lab: Evaluating acme-assist
This lab uses the ai-lab kit. Everything runs on the free local model path.
Reviewing the Quick Reference and Common Mistakes
Quick reference Common mistakes Skipping the golden dataset and relying on "it seemed to work" feels fast, but it...
Career Impact
Roles that use the skills in this module.
AI Engineer
MLOps Engineer
Platform Engineer
DevOps Engineer
Next Modules
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
Twenty to thirty well chosen cases is enough to start. Cover common questions, edge cases, questions with no answer, and the mistakes that would hurt most. Then keep adding real failures from production, because the dataset is never finished.
It can, but treat its scores as an estimate, not a fact. Use a pinned judge model, ask for validated JSON, and read a sample of judged cases by hand every few weeks. Use plain rules wherever a rule is enough.
A regression is a drop caused by a change you made, such as a new prompt or model. You catch it by re-running the golden dataset before you merge. Drift is a slow decline with no code change, caught by watching production traffic.
Yes, a small one. Three or four thresholds in a JSON file that your script checks costs almost nothing. It turns 'does this look okay' into a pass or fail that nobody has to argue about at midnight.