Skip to main content

Linux Internals for SRE

Master Linux kernel internals for SRE work - process lifecycle, memory management, cgroups, namespaces, CPU scheduler, and production diagnosis using /proc, strace, and eBPF.

Prerequisites
~55 minutes
17 Topics
Hands-on Scenarios

What You'll Learn

The Production Incident That Changes How You Think About Linux

It is 3 AM. Hotstar's streaming service is throwing 503s for 2 million concurrent users during an IPL match. Your on-call rotation fires.

Understanding the Linux Boot Process and Failure Points

Before a single process runs, the kernel has to get off the ground.

Process Lifecycle at the Syscall Level

You already know processes exist. This section explains how they are born, how they communicate, and how they die at the kernel level - because most...

File Descriptors and the Open File Table

This is the concept behind the Hotstar incident from the introduction.

Memory Management - What the Kernel Does With Your RAM

Running out of memory is one of the most misdiagnosed production problems. People see low free memory and panic.

Reading /proc and /sys Without External Tools

During a production incident, tools like ps, top, and htop may not be installed (inside a minimal container), may be broken (their own processes...

Skills You'll Master

LINUXKERNELSRECGROUPSOBSERVABILITY

Curriculum Index17 topics

1

The Production Incident That Changes How You Think About Linux

It is 3 AM. Hotstar's streaming service is throwing 503s for 2 million concurrent users during an IPL match.

2

Understanding the Linux Boot Process and Failure Points

Before a single process runs, the kernel has to get off the ground.

3

Process Lifecycle at the Syscall Level

You already know processes exist. This section explains how they are born, how they communicate, and how they die at...

4

File Descriptors and the Open File Table

This is the concept behind the Hotstar incident from the introduction.

5

Memory Management - What the Kernel Does With Your RAM

Running out of memory is one of the most misdiagnosed production problems. People see low free memory and panic.

6

Reading /proc and /sys Without External Tools

During a production incident, tools like ps, top, and htop may not be installed (inside a minimal container), may be...

7

cgroups - The Kernel Mechanism Behind Every Container

When you run kubectl apply -f deployment.yaml and set resources.limits.memory: 512Mi, something real happens in the...

8

Linux Namespaces - What a Container Really Is

A container is not a separate kernel. It is not a VM.

9

The CPU Scheduler - Who Gets CPU Time and When

The Linux CPU scheduler makes thousands of decisions per second about which process runs on which CPU core.

10

I/O Analysis - Finding Disk Bottlenecks

Disk I/O problems are one of the hardest to diagnose because they manifest as everything else - high CPU iowait, slow...

11

Network Stack Internals - Below the Application Layer

Network problems that look like application problems - connections timing out, intermittent failures under load, high...

12

Process Groups, Sessions, and Why Ctrl+C Works

This section explains something most engineers use every day but rarely think about: why pressing Ctrl+C kills your...

13

The USE Method - A Systematic Diagnostic Framework

With all the tools and internals covered above, you now need a mental framework for using them systematically.

14

Flame Graphs and eBPF - Finding the Real Bottleneck

When USE tells you CPU utilisation is high but you do not know WHY - which function, which system call, which code path...

15

Hands-on Lab

This lab takes you through five real SRE diagnostic scenarios using kernel-level tools.

16

Quick Reference

Essential /proc paths Key one-liners for incidents bpftrace one-liners

17

Quick Reference and Common Mistakes

Treating high iowait as a disk problem is one of the most expensive diagnostic mistakes in SRE work.

Career Impact

Roles that use the skills in this module.

  • Site Reliability Engineer

  • DevOps Engineer

  • Platform Engineer

  • Cloud Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet