Linux Kernel
The Linux kernel is the core of the operating system — the software that manages hardware, memory, processes, and system calls. It runs in privileged kernel space while applications run in user space. The kernel is the layer that makes all higher-level Linux functionality possible.
Understanding the Linux Kernel
What Is the Linux Kernel in Simple Terms
The kernel is the operating system's core. Everything else — bash, nginx, Python, Docker, even the desktop — runs on top of the kernel. The kernel manages physical hardware (CPU, memory, disk, network), schedules which processes run when, and provides controlled access to hardware through system calls.
As a DevOps engineer, you interact with the kernel constantly — through sysctl parameters, kernel modules, system calls, and the /proc and /sys virtual filesystems — even if you never touch the kernel source code.
How It Works
+------------------------------------------+| User Space || nginx, postgres, node, bash, docker |+------------------------------------------+ | system calls (read, write, open) v+------------------------------------------+| Kernel Space || Process scheduler, memory manager || VFS (virtual filesystem layer) || Network stack (TCP/IP) || Device drivers |+------------------------------------------+ | hardware access v+------------------------------------------+| Hardware || CPU, RAM, Disk, Network interface card |+------------------------------------------+Practical Commands
## Check kernel versionuname -r## 5.15.0-1031-aws <- kernel version, build, and platform uname -a## Linux mumbai-prod-node-1 5.15.0-1031-aws #35-Ubuntu SMP ... ## View kernel messages (hardware events, OOM kills, driver messages)dmesg | tail -20dmesg | grep -i 'error\|fail\|oom' ## View kernel messages with timestampsdmesg -T | tail -20 ## Kernel parameters at runtimesysctl -a | grep net.ipv4sysctl net.ipv4.tcp_syncookies ## view one parametersudo sysctl -w net.ipv4.tcp_syncookies=1 ## set immediately ## Persist sysctl settingssudo tee /etc/sysctl.d/99-custom.conf << 'EOF'net.ipv4.tcp_syncookies = 1vm.swappiness = 10EOFsudo sysctl -p /etc/sysctl.d/99-custom.conf ## List loaded kernel moduleslsmod ## Load a kernel modulesudo modprobe br_netfilter ## needed for Kubernetes networking ## Check if module is loadedlsmod | grep br_netfilter ## Make module load at bootecho 'br_netfilter' | sudo tee /etc/modules-load.d/k8s.conf ## View kernel boot parameterscat /proc/cmdline ## Available memory and swapcat /proc/meminfo | grep -E 'MemTotal|MemAvailable|SwapTotal'Troubleshooting
| Symptom | Command | What to Check |
|---|---|---|
| OOM kills happening | `dmesg | grep -i oom` |
| Kernel module missing | `lsmod | grep modulename` |
| High system CPU (sy) | perf top |
Kernel code consuming CPU |
| Network issues | sysctl net.ipv4 |
Kernel network parameters |
Tip
dmesg -T | grep -i oomis your first command when applications are dying unexpectedly. The OOM (Out of Memory) killer logs every process it kills, including why — the process name, PID, and memory usage. This immediately tells you if memory pressure is causing your application crashes.
Remembersysctl changes made with
sysctl -ware immediate but temporary — lost on reboot. Always write them to/etc/sysctl.d/99-custom.confand runsysctl -pto make them persistent. Kubernetes setup guides often require specific sysctl values (likenet.bridge.bridge-nf-call-iptables=1) that must persist across reboots.
Frequently Asked Questions
How does the kernel's separation of kernel space and user space actually get enforced at the hardware level?
The CPU itself provides privilege rings (kernel space runs in ring 0 with unrestricted hardware access, user-space applications run in ring 3 with restricted access), and any interaction between them — reading a file, opening a network socket — happens through system calls, a defined interface where the CPU switches privilege levels in a controlled way. This is what stops a bug in a normal application from directly corrupting kernel memory or arbitrary hardware state.
Why does a kernel panic differ from an application crash in terms of recovery?
An application crash is contained — the OS reclaims its memory and other processes continue unaffected. A kernel panic means the kernel itself hit an unrecoverable inconsistency (corrupted internal state, a failed hardware assumption) and halts the entire system, because there's no higher authority left to safely intervene; the machine typically requires a hard reboot. This is why kernel-level code (drivers, modules) is held to a much higher reliability bar than userspace code — a bug there can take down everything, not just one process.