Skip to main content

Linux Kernel

The Linux kernel is the core of the operating system — the software that manages hardware, memory, processes, and system calls. It runs in privileged kernel space while applications run in user space. The kernel is the layer that makes all higher-level Linux functionality possible.

Understanding the Linux Kernel

What Is the Linux Kernel in Simple Terms

The kernel is the operating system's core. Everything else — bash, nginx, Python, Docker, even the desktop — runs on top of the kernel. The kernel manages physical hardware (CPU, memory, disk, network), schedules which processes run when, and provides controlled access to hardware through system calls.

As a DevOps engineer, you interact with the kernel constantly — through sysctl parameters, kernel modules, system calls, and the /proc and /sys virtual filesystems — even if you never touch the kernel source code.

How It Works

Bash
+------------------------------------------+
| User Space |
| nginx, postgres, node, bash, docker |
+------------------------------------------+
| system calls (read, write, open)
v
+------------------------------------------+
| Kernel Space |
| Process scheduler, memory manager |
| VFS (virtual filesystem layer) |
| Network stack (TCP/IP) |
| Device drivers |
+------------------------------------------+
| hardware access
v
+------------------------------------------+
| Hardware |
| CPU, RAM, Disk, Network interface card |
+------------------------------------------+

Practical Commands

Bash
## Check kernel version
uname -r
## 5.15.0-1031-aws <- kernel version, build, and platform
uname -a
## Linux mumbai-prod-node-1 5.15.0-1031-aws #35-Ubuntu SMP ...
## View kernel messages (hardware events, OOM kills, driver messages)
dmesg | tail -20
dmesg | grep -i 'error\|fail\|oom'
## View kernel messages with timestamps
dmesg -T | tail -20
## Kernel parameters at runtime
sysctl -a | grep net.ipv4
sysctl net.ipv4.tcp_syncookies ## view one parameter
sudo sysctl -w net.ipv4.tcp_syncookies=1 ## set immediately
## Persist sysctl settings
sudo tee /etc/sysctl.d/99-custom.conf << 'EOF'
net.ipv4.tcp_syncookies = 1
vm.swappiness = 10
EOF
sudo sysctl -p /etc/sysctl.d/99-custom.conf
## List loaded kernel modules
lsmod
## Load a kernel module
sudo modprobe br_netfilter ## needed for Kubernetes networking
## Check if module is loaded
lsmod | grep br_netfilter
## Make module load at boot
echo 'br_netfilter' | sudo tee /etc/modules-load.d/k8s.conf
## View kernel boot parameters
cat /proc/cmdline
## Available memory and swap
cat /proc/meminfo | grep -E 'MemTotal|MemAvailable|SwapTotal'

Troubleshooting

Symptom Command What to Check
OOM kills happening `dmesg grep -i oom`
Kernel module missing `lsmod grep modulename`
High system CPU (sy) perf top Kernel code consuming CPU
Network issues sysctl net.ipv4 Kernel network parameters
Tip

dmesg -T | grep -i oom is your first command when applications are dying unexpectedly. The OOM (Out of Memory) killer logs every process it kills, including why — the process name, PID, and memory usage. This immediately tells you if memory pressure is causing your application crashes.

Remember

sysctl changes made with sysctl -w are immediate but temporary — lost on reboot. Always write them to /etc/sysctl.d/99-custom.conf and run sysctl -p to make them persistent. Kubernetes setup guides often require specific sysctl values (like net.bridge.bridge-nf-call-iptables=1) that must persist across reboots.

Frequently Asked Questions

How does the kernel's separation of kernel space and user space actually get enforced at the hardware level?

The CPU itself provides privilege rings (kernel space runs in ring 0 with unrestricted hardware access, user-space applications run in ring 3 with restricted access), and any interaction between them — reading a file, opening a network socket — happens through system calls, a defined interface where the CPU switches privilege levels in a controlled way. This is what stops a bug in a normal application from directly corrupting kernel memory or arbitrary hardware state.

Why does a kernel panic differ from an application crash in terms of recovery?

An application crash is contained — the OS reclaims its memory and other processes continue unaffected. A kernel panic means the kernel itself hit an unrecoverable inconsistency (corrupted internal state, a failed hardware assumption) and halts the entire system, because there's no higher authority left to safely intervene; the machine typically requires a hard reboot. This is why kernel-level code (drivers, modules) is held to a much higher reliability bar than userspace code — a bug there can take down everything, not just one process.