How AI Agents Work and Tool Calling
Learn how ops agents work: the reason-act-observe loop, tool definitions, workflows versus agents, and building an investigator with LangGraph.
What You'll Learn
Understanding What an Ops Agent Is
An ops agent is a language model, a set of tools, and a loop, and nothing more mysterious than that.
Tool Calling and the Execution Boundary
Tool calling is how a model asks for a function to be run: it replies with a structured request instead of prose.
Writing Tool Definitions as Contracts
A tool definition is a contract between you and the model, and its quality sets the ceiling on how well the agent behaves.
Choosing Between a Workflow and an Agent
Most ops automation should be a fixed workflow with one model judgement inside it, and an agent is the right choice only when the next step truly...
Running Independent Tools in Parallel
When two tool calls do not depend on each other, run them at the same time.
Building the Incident Investigator Loop
The investigator is a short loop over real tools running against aiops-lab, with a hard step limit and a log of every call.
Skills You'll Master
Curriculum Index9 topics
Understanding What an Ops Agent Is
An ops agent is a language model, a set of tools, and a loop, and nothing more mysterious than that.
Tool Calling and the Execution Boundary
Tool calling is how a model asks for a function to be run: it replies with a structured request instead of prose.
Writing Tool Definitions as Contracts
A tool definition is a contract between you and the model, and its quality sets the ceiling on how well the agent...
Choosing Between a Workflow and an Agent
Most ops automation should be a fixed workflow with one model judgement inside it, and an agent is the right choice...
Running Independent Tools in Parallel
When two tool calls do not depend on each other, run them at the same time.
Building the Incident Investigator Loop
The investigator is a short loop over real tools running against aiops-lab, with a hard step limit and a log of every...
Rebuilding It with LangGraph
LangGraph expresses the same loop as an explicit graph of nodes and edges.
Running the Hands-on Lab
This lab faults payment-service, runs the investigator in plain Python and in LangGraph, and compares the two traces.
Quick Reference, Common Mistakes, and Interview Questions
Quick reference Common mistakes Building an agent where a workflow would do is the most common one.
Career Impact
Roles that use the skills in this module.
- High Demand
AIOps Engineer
$130k Average a year
- Very High Demand
ML Engineer
$140k Average a year
- High Demand
Site Reliability Engineer
$145k Average a year
- Growing
Platform Engineer
$135k Average a year
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
An LLM turns text into text and knows nothing about your systems right now. An agent wraps the LLM in a loop with tools, so it can fetch live data, look at the result, and decide what to do next. The model supplies reasoning and your code supplies the actions.
No. The model only proposes a tool call as structured text. Your code validates the arguments, runs the function, and sends the result back. That gap is where limits, logging, and approvals live.
Use a fixed workflow when you already know the steps, which is true for most ops automation. Use an agent when the next step really depends on what the last step found. Workflows are cheaper, faster, and easier to test.
A confused agent can keep calling tools forever, burning time and money. A hard cap on steps makes a stuck run stop and report instead. Pair it with logging so you can see what it was doing.