Skip to main content

AIOps Platforms and the AI SRE Pattern

Learn what commercial AIOps platforms automate, when the open-source stack is enough, and how to judge an AI investigation by speed and accuracy.

~2.5 hours
7 Topics
Hands-on Scenarios

What You'll Learn

Understanding What an AIOps Platform Adds

An AIOps platform is a product that ingests operational data, learns what normal looks like, groups related problems together, and points at a likely...

Comparing Three Platform Approaches at the Concept Level

Datadog Watchdog, Dynatrace, and Splunk IT Service Intelligence (ITSI) all fall under the AIOps label, but each documents a different centre of...

Building Versus Buying: When Open Source Is Enough

You should build with the open-source stack when your team is small, your service count is modest, and the glue work is something you can maintain.

Understanding the AI SRE Pattern

The AI SRE pattern is a way of using AI in on-call: when an alert fires, an AI system investigates first using read-only tools, and hands the human...

Judging an AI Investigation with a Scorecard

You judge an AI investigation by comparing its summary with what the human investigation eventually found, using the same few criteria every time.

Measuring Before and After Adoption

You measure a tool's value by recording how long incidents took to detect and resolve before adopting it, then again after.

Skills You'll Master

AIOPSAIOPS-PLATFORMSAI-SREINCIDENT-RESPONSEOBSERVABILITY

Curriculum Index7 topics

Career Impact

Roles that use the skills in this module.

  • AI Engineer

  • MLOps Engineer

  • Platform Engineer

  • Site Reliability Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

A platform bundles several jobs into one product: collecting data, learning what is normal, grouping related alerts, using a map of service dependencies to point at a likely cause, and presenting it all in one console. You can build each piece with open-source tools, but you own the glue, the tuning, and the on-call burden of the glue itself. A platform trades money and lock-in for less assembly work.

Not quite. An AIOps platform mostly detects and groups problems. An AI SRE is a pattern where an AI system investigates an alert first, using read-only tools, and gives the on-call human a summary with evidence. Some platforms now offer this, and you can also build it yourself later in this roadmap.

Compare its summary with what the human investigation finally found. Check whether it named the real cause, cited evidence you can verify, separated the cause from downstream victims, and was honest about uncertainty. A scorecard with a few fixed criteria makes this repeatable instead of a matter of opinion.

No. The open-source stack in this roadmap covers detection, correlation, investigation, and measurement for a small or mid-sized system. A platform starts to make sense when the glue work, the number of services, or the data volume outgrows what your team can maintain.