Anomaly Detection for Ops: Baselines and Isolation Forest
Learn to catch ops anomalies early: rolling and seasonal baselines first, then Isolation Forest, with precision and recall proving fewer false pages.
What You'll Learn
Understanding Anomaly Detection and Why It Matters in Ops
You are the SRE building acme-shop's AIOps capability. Earlier modules gave you metrics, logs, and traces for cart-service and payment-service.
Understanding Supervised vs Unsupervised Anomaly Detection
Before choosing a method, it helps to know why ops work almost always uses methods that do not need labels, and what the simple statistical methods...
Building a Seasonal Baseline That Isolation Forest Must Beat
A model has to earn its complexity. The fair test is not a global average.
Understanding Isolation Forest - The Core Idea
Isolation Forest is one of the most used anomaly detection algorithms.
How the Isolation Forest Algorithm Works Step by Step
Knowing the intuition is not enough to use it well. This topic covers the forest, the one setting that matters most, and what the outputs mean.
Setting Up and Running Isolation Forest in Python
Here is the simplest possible run: raw metrics in, labels out. The next topics show why it is not enough.
Skills You'll Master
Curriculum Index10 topics
Understanding Anomaly Detection and Why It Matters in Ops
You are the SRE building acme-shop's AIOps capability.
Understanding Supervised vs Unsupervised Anomaly Detection
Before choosing a method, it helps to know why ops work almost always uses methods that do not need labels, and what...
Building a Seasonal Baseline That Isolation Forest Must Beat
A model has to earn its complexity. The fair test is not a global average.
Understanding Isolation Forest - The Core Idea
Isolation Forest is one of the most used anomaly detection algorithms.
How the Isolation Forest Algorithm Works Step by Step
Knowing the intuition is not enough to use it well.
Setting Up and Running Isolation Forest in Python
Here is the simplest possible run: raw metrics in, labels out. The next topics show why it is not enough.
Evaluating Your Anomaly Detection Results
A model that flags everything is useless, and so is one that flags nothing.
Applying Isolation Forest to Real Ops Metrics
The jump from a tutorial to real metrics is mostly about giving the model context.
Hands-on Lab - Build a Complete Anomaly Detector
📌 Remember: This lab is free and runs on your laptop in about 20 minutes. It needs Python 3.12 and a terminal.
Quick Reference, Limitations, and Interview Questions
When to Use Isolation Forest It is a strong default when you have unlabelled metrics, several features to judge...
Career Impact
Roles that use the skills in this module.
- High Demand
AIOps Engineer
$130k Average a year
- High Demand
Site Reliability Engineer
$145k Average a year
- Very High Demand
ML Engineer
$140k Average a year
- Growing
Platform Engineer
$135k Average a year
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
A fixed threshold cannot know that 60 percent CPU is normal at noon and strange at 3am. Thresholds are either too sensitive and cause alert fatigue, or too loose and miss slow degradation. Baselines and models judge a value against what is normal for that time.
No. It is unsupervised: it learns the shape of normal data and flags points that are easy to separate from it. Labels are only needed if you want to measure how well it works.
Contamination is the share of data you expect to be anomalous. It sets the score cut-off, so a value that is too high floods you with false pages and one that is too low misses incidents.
When incidents show up clearly in one metric, a seasonal baseline is simpler, faster, and easier to explain. Use Isolation Forest only if it beats the baseline on the same data.