AI Agents, Workflows, and Reliability Engineering
Learn to build AI agents that act safely: tool design, explicit state, LangGraph, MCP, retries, idempotency, step limits, and human approval gates.
What You'll Learn
Understanding Why Agent Reliability Matters as Much as Agent Intelligence
acme-assist can already answer from the help centre. Now it has to do something harder: look up an order and issue a refund.
Designing Tools for Reliable Function Calling
An agent is only as reliable as its tools.
Designing Agent State Before You Orchestrate
Tools come first, then state, then orchestration.
Building Agent Orchestration With LangGraph
With tools and state defined, you need a way to run the loop with explicit control over transitions and stopping.
Connecting Agents to Tools With the Model Context Protocol
So far tools have been functions inside your codebase.
Engineering Reliability: Retries, Idempotency, and Timeouts
Good tool design and state make an agent capable and inspectable.
Skills You'll Master
Curriculum Index13 topics
Understanding Why Agent Reliability Matters as Much as Agent Intelligence
acme-assist can already answer from the help centre.
Designing Tools for Reliable Function Calling
An agent is only as reliable as its tools.
Designing Agent State Before You Orchestrate
Tools come first, then state, then orchestration.
Building Agent Orchestration With LangGraph
With tools and state defined, you need a way to run the loop with explicit control over transitions and stopping.
Connecting Agents to Tools With the Model Context Protocol
So far tools have been functions inside your codebase.
Engineering Reliability: Retries, Idempotency, and Timeouts
Good tool design and state make an agent capable and inspectable.
Preventing Runaway Agents
Retries and idempotency guard one action.
Adding Checkpointing and Human Approval
An agent that fails partway through should not restart from nothing, and an agent about to do something hard to reverse...
Logging and Debugging Agent Behavior
Every guardrail above assumes you can see what the agent did.
Hands-On Lab: Building a Reliable Refund Agent
The agent works on a mock payments service for acme-shop, so nothing real is touched.
Quick Reference
A compact map of each guardrail and the failure it prevents.
Common Mistakes
Overlapping tools with similar descriptions make the model pick between them inconsistently, because nothing says when...
Reviewing What You Built and What Comes Next
You started with an agent that could refund the same transaction three times, and you finished with one that cannot.
Career Impact
Roles that use the skills in this module.
AI Engineer
MLOps Engineer
Platform Engineer
Site Reliability Engineer
Next Modules
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
A workflow follows steps you fixed in code, and the model only handles judgment calls inside them. An agent chooses its own next action at runtime. Workflows are easier to test and predict, so use an agent only when the sequence of steps cannot be known in advance.
A model that keeps getting an error often retries a small variation of the same call, hoping the next attempt works. Nothing stops it unless you add a hard step limit and loop detection around the agent. These are control guardrails, not model improvements.
It is a unique ID attached to one logical operation, such as one refund. If the same key arrives twice, the receiving service returns the original result instead of repeating the action. You generate it once and reuse it on every retry.
No. MCP standardizes how an agent connects to tools, but it does not validate arguments, authorize calls, or prevent duplicates. You still need every guardrail you would put around a local tool.