Skip to main content

Anomaly Detection for Ops: Baselines and Isolation Forest

Learn to catch ops anomalies early: rolling and seasonal baselines first, then Isolation Forest, with precision and recall proving fewer false pages.

~3 hours
10 Topics
Hands-on Scenarios

What You'll Learn

Understanding Anomaly Detection and Why It Matters in Ops

You are the SRE building acme-shop's AIOps capability. Earlier modules gave you metrics, logs, and traces for cart-service and payment-service.

Understanding Supervised vs Unsupervised Anomaly Detection

Before choosing a method, it helps to know why ops work almost always uses methods that do not need labels, and what the simple statistical methods...

Building a Seasonal Baseline That Isolation Forest Must Beat

A model has to earn its complexity. The fair test is not a global average.

Understanding Isolation Forest - The Core Idea

Isolation Forest is one of the most used anomaly detection algorithms.

How the Isolation Forest Algorithm Works Step by Step

Knowing the intuition is not enough to use it well. This topic covers the forest, the one setting that matters most, and what the outputs mean.

Setting Up and Running Isolation Forest in Python

Here is the simplest possible run: raw metrics in, labels out. The next topics show why it is not enough.

Skills You'll Master

ANOMALY-DETECTIONISOLATION-FORESTSCIKIT-LEARNAIOPSALERTING

Curriculum Index10 topics

Career Impact

Roles that use the skills in this module.

  • AIOps Engineer

    $130k Average a year

    High Demand
  • Site Reliability Engineer

    $145k Average a year

    High Demand
  • ML Engineer

    $140k Average a year

    Very High Demand
  • Platform Engineer

    $135k Average a year

    Growing
See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

A fixed threshold cannot know that 60 percent CPU is normal at noon and strange at 3am. Thresholds are either too sensitive and cause alert fatigue, or too loose and miss slow degradation. Baselines and models judge a value against what is normal for that time.

No. It is unsupervised: it learns the shape of normal data and flags points that are easy to separate from it. Labels are only needed if you want to measure how well it works.

Contamination is the share of data you expect to be anomalous. It sets the score cut-off, so a value that is too high floods you with false pages and one that is too low misses incidents.

When incidents show up clearly in one metric, a seasonal baseline is simpler, faster, and easier to explain. Use Isolation Forest only if it beats the baseline on the same data.