A production server at Hotstar handling IPL traffic is running hundreds of processes simultaneously — nginx workers, application threads, log shippers, health checks, and cron jobs. When something goes wrong — a runaway process consuming 100% CPU, a service that will not restart, a disk that fills up silently — the tools in this pillar are what diagnose and fix it in minutes rather than hours.
What This Pillar Covers
- Inspecting and controlling processes with ps, top, htop, kill, and Linux signals
- Managing production services with systemd — unit files, journalctl, and restart policies
- Diagnosing CPU, memory, disk I/O, and network bottlenecks with vmstat, iostat, and free
- Scheduling automated tasks with cron and systemd timers with proper logging and error handling
Who This Is For
DevOps engineers and site reliability engineers responsible for diagnosing performance issues and keeping production Linux servers healthy under load.
Why This Matters in Production
During peak trading hours at Zerodha, a runaway Java process consumed all available CPU, making the trading platform unresponsive. Engineers who knew how to identify the process with htop, check its open file descriptors, and send the correct signal recovered the server in under 3 minutes.
Prerequisites
- Linux Fundamentals — filesystem navigation and basic commands
- Understanding of what a program and a process are
- Basic familiarity with system services (web server, database)