AI Safety and Responsible AI
Learn to defend AI products before launch: prompt injection, retrieval authorization, moderation, bias testing, privacy, and a risk-sized checklist.
What You'll Learn
Understanding Why acme-assist Is Not Ready to Launch
The playbook that reached a customer acme-assist can now answer help-centre questions with citations, and the support team loves the demo.
Understanding Prompt Injection
The attack that looks like an ordinary question A customer asks acme-assist something normal: "How long does a UPI refund take?" Nothing about the...
Separating Trusted Instructions from Untrusted Content
Why the role of a message matters Chat models treat message roles differently.
Authorizing Retrieval Before the Model Sees Anything
Relevance is not authorization A vector search answers one question: which chunks are closest in meaning?
Screening Outputs and Authorizing Actions
Checking what the model is about to say Output screening checks the final answer before the customer sees it.
Adding a Real Moderation Layer
Moderation is not injection defence A moderation layer is a separate check that screens inputs and outputs for harmful or policy-violating content...
Skills You'll Master
Curriculum Index12 topics
Understanding Why acme-assist Is Not Ready to Launch
The playbook that reached a customer acme-assist can now answer help-centre questions with citations, and the support...
Understanding Prompt Injection
The attack that looks like an ordinary question A customer asks acme-assist something normal: "How long does a UPI...
Separating Trusted Instructions from Untrusted Content
Why the role of a message matters Chat models treat message roles differently.
Authorizing Retrieval Before the Model Sees Anything
Relevance is not authorization A vector search answers one question: which chunks are closest in meaning?
Screening Outputs and Authorizing Actions
Checking what the model is about to say Output screening checks the final answer before the customer sees it.
Adding a Real Moderation Layer
Moderation is not injection defence A moderation layer is a separate check that screens inputs and outputs for harmful...
Testing for Bias and Fairness
Why the same question can get different answers Bias means a system produces systematically different outcomes for...
Protecting Privacy and Personal Data
What leaves your infrastructure Every call to a hosted model sends the prompt, including retrieved text, to a third...
Sizing Guardrails to Risk and Logging Safely
Matching intensity to risk An internal tool used by twelve engineers does not need the guardrails of a public chatbot...
Running the Safety Lab on acme-assist
Cost: everything here runs free on local models.
What You Built and What Comes Next
You attacked acme-assist with a poisoned help article, measured the attack, added structural separation and output...
Quick Reference and Common Mistakes
Putting retrieved text in the system message happens because it feels like background instruction.
Career Impact
Roles that use the skills in this module.
AI Engineer
MLOps Engineer
Platform Engineer
Next Modules
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding Sheet