Skip to main content

How AI Agents Work and Tool Calling

Learn how ops agents work: the reason-act-observe loop, tool definitions, workflows versus agents, and building an investigator with LangGraph.

~5 hours
9 Topics
Hands-on Scenarios

What You'll Learn

Understanding What an Ops Agent Is

An ops agent is a language model, a set of tools, and a loop, and nothing more mysterious than that.

Tool Calling and the Execution Boundary

Tool calling is how a model asks for a function to be run: it replies with a structured request instead of prose.

Writing Tool Definitions as Contracts

A tool definition is a contract between you and the model, and its quality sets the ceiling on how well the agent behaves.

Choosing Between a Workflow and an Agent

Most ops automation should be a fixed workflow with one model judgement inside it, and an agent is the right choice only when the next step truly...

Running Independent Tools in Parallel

When two tool calls do not depend on each other, run them at the same time.

Building the Incident Investigator Loop

The investigator is a short loop over real tools running against aiops-lab, with a hard step limit and a log of every call.

Skills You'll Master

AI-AGENTSTOOL-CALLINGLANGGRAPHREACT-PATTERNAIOPS

Curriculum Index9 topics

Career Impact

Roles that use the skills in this module.

  • AIOps Engineer

    $130k Average a year

    High Demand
  • ML Engineer

    $140k Average a year

    Very High Demand
  • Site Reliability Engineer

    $145k Average a year

    High Demand
  • Platform Engineer

    $135k Average a year

    Growing
See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

An LLM turns text into text and knows nothing about your systems right now. An agent wraps the LLM in a loop with tools, so it can fetch live data, look at the result, and decide what to do next. The model supplies reasoning and your code supplies the actions.

No. The model only proposes a tool call as structured text. Your code validates the arguments, runs the function, and sends the result back. That gap is where limits, logging, and approvals live.

Use a fixed workflow when you already know the steps, which is true for most ops automation. Use an agent when the next step really depends on what the last step found. Workflows are cheaper, faster, and easier to test.

A confused agent can keep calling tools forever, burning time and money. A hard cap on steps makes a stuck run stop and report instead. Pair it with logging so you can see what it was doing.