AIOps Platforms and the AI SRE Pattern
Learn what commercial AIOps platforms automate, when the open-source stack is enough, and how to judge an AI investigation by speed and accuracy.
What You'll Learn
Understanding What an AIOps Platform Adds
An AIOps platform is a product that ingests operational data, learns what normal looks like, groups related problems together, and points at a likely...
Comparing Three Platform Approaches at the Concept Level
Datadog Watchdog, Dynatrace, and Splunk IT Service Intelligence (ITSI) all fall under the AIOps label, but each documents a different centre of...
Building Versus Buying: When Open Source Is Enough
You should build with the open-source stack when your team is small, your service count is modest, and the glue work is something you can maintain.
Understanding the AI SRE Pattern
The AI SRE pattern is a way of using AI in on-call: when an alert fires, an AI system investigates first using read-only tools, and hands the human...
Judging an AI Investigation with a Scorecard
You judge an AI investigation by comparing its summary with what the human investigation eventually found, using the same few criteria every time.
Measuring Before and After Adoption
You measure a tool's value by recording how long incidents took to detect and resolve before adopting it, then again after.
Skills You'll Master
Curriculum Index7 topics
Understanding What an AIOps Platform Adds
An AIOps platform is a product that ingests operational data, learns what normal looks like, groups related problems...
Comparing Three Platform Approaches at the Concept Level
Datadog Watchdog, Dynatrace, and Splunk IT Service Intelligence (ITSI) all fall under the AIOps label, but each...
Building Versus Buying: When Open Source Is Enough
You should build with the open-source stack when your team is small, your service count is modest, and the glue work is...
Understanding the AI SRE Pattern
The AI SRE pattern is a way of using AI in on-call: when an alert fires, an AI system investigates first using...
Judging an AI Investigation with a Scorecard
You judge an AI investigation by comparing its summary with what the human investigation eventually found, using the...
Measuring Before and After Adoption
You measure a tool's value by recording how long incidents took to detect and resolve before adopting it, then again...
Working Through the Platform Evaluation Exercise
This exercise needs no cluster and no vendor account.
Career Impact
Roles that use the skills in this module.
AI Engineer
MLOps Engineer
Platform Engineer
Site Reliability Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
A platform bundles several jobs into one product: collecting data, learning what is normal, grouping related alerts, using a map of service dependencies to point at a likely cause, and presenting it all in one console. You can build each piece with open-source tools, but you own the glue, the tuning, and the on-call burden of the glue itself. A platform trades money and lock-in for less assembly work.
Not quite. An AIOps platform mostly detects and groups problems. An AI SRE is a pattern where an AI system investigates an alert first, using read-only tools, and gives the on-call human a summary with evidence. Some platforms now offer this, and you can also build it yourself later in this roadmap.
Compare its summary with what the human investigation finally found. Check whether it named the real cause, cited evidence you can verify, separated the cause from downstream victims, and was honest about uncertainty. A scorecard with a few fixed criteria makes this repeatable instead of a matter of opinion.
No. The open-source stack in this roadmap covers detection, correlation, investigation, and measurement for a small or mid-sized system. A platform starts to make sense when the glue work, the number of services, or the data volume outgrows what your team can maintain.