Bash and Linux

Find the Bottleneck

mediumSystem monitoring

Problem statement

Read vmstat samples from three slow servers and decide for each one whether it is CPU-bound, short of memory (swapping) or waiting on disk. "The server is slow" has very different fixes depending on which resource is the bottleneck, and vmstat tells you in one screen.

Each file holds vmstat 1 3 output: two header lines, then three samples taken one second apart.

cpu.txt

TEXT
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
1 0 0 812000 50000 900000 0 0 5 10 200 300 5 2 92 1 0
8 0 0 810000 50000 900000 0 0 0 4 3000 5000 88 10 2 0 0
9 0 0 809000 50000 900000 0 0 0 8 3100 5200 90 9 1 0 0

io.txt

TEXT
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
1 0 0 812000 50000 900000 0 0 5 10 200 300 5 2 92 1 0
1 4 0 790000 50000 910000 0 0 9000 12000 900 1200 6 4 45 45 0
0 5 0 785000 50000 915000 0 0 9500 11000 950 1300 5 3 42 50 0

mem.txt

TEXT
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
1 0 400000 90000 2000 60000 5 10 50 80 300 400 10 5 80 5 0
3 1 520000 30000 1000 40000 300 500 900 1200 1500 2500 20 10 40 30 0
4 2 600000 25000 1000 35000 450 600 1100 1500 1700 2700 22 12 30 36 0

For each file, ignore the first sample (it is an average since boot), average the other two, and print the CPU use (us + sy), I/O wait (wa) and swap traffic (si + so), followed by a verdict.

Expected output:

โ—ˆ DIAGRAM
cpu.txt cpu 98% wait 0% swap 0 KiB/s -> CPU-bound
io.txt cpu 9% wait 48% swap 0 KiB/s -> I/O-bound (waiting on disk)
mem.txt cpu 32% wait 33% swap 925 KiB/s -> memory-bound (swapping)

Hints

Hint 1: Skip the two header lines and the first sample with NR > 3. The columns are r b swpd free buff cache si so bi bo in cs us sy id wa st, so si is field 7, so field 8, us 13, sy 14 and wa 16.

Approach

Optimal: vmstat columns and rules

Covers: vmstat columns, why the first line differs, us sy id wa st, si so, r and b, awk sums and averages, rule order, FILENAME.

Three resources, three signs. A slow Linux machine is almost always waiting on one of three things. vmstat shows all of them in one line per second:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart LR subgraph CPU["CPU"] direction TB C1["us + sy high
r above CPUs"]:::red end subgraph MEM["Memory"] direction TB M1["si, so above 0
free very low"]:::yellow end subgraph IO["Disk"] direction TB I1["wa high
b above 0"]:::blue end CPU ~~~ MEM ~~~ IO classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px style CPU fill:transparent,stroke:#dc2626,stroke-width:2px style MEM fill:transparent,stroke:#d97706,stroke-width:2px style IO fill:transparent,stroke:#2563eb,stroke-width:2px
Column Group Means Worry when
r procs tasks running or waiting for a CPU above the CPU count
b procs tasks blocked, waiting on I/O above 0 for a while
si, so swap KiB per second read from / written to swap above 0 for a while
bi, bo io blocks read from / written to disk depends on the disk
us, sy cpu % time in programs / in the kernel together above 80
id cpu % idle near 0
wa cpu % idle while waiting for disk above 20
st cpu % stolen by the host on a VM above a few %

The first line lies. The first sample vmstat prints is an average since the machine booted, which can be weeks. Always skip it and look at the following lines. That is why the hint uses NR > 3: two header lines plus that first sample.

Why the rule order matters. When a machine runs out of RAM, it swaps, and swapping is disk I/O, so wa rises too. If you checked wa first, you would blame the disk when the real cause is memory. So the rules run in this order:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB S{{"si + so above 0?"}}:::yellow --> SM["memory-bound"]:::yellow S --> W{{"wa 20% or more?"}}:::blue W --> WM["I/O-bound"]:::blue W --> C{{"us + sy 80% or more?"}}:::red C --> CM["CPU-bound"]:::red C --> H["healthy"]:::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px

Walking through the code. The # Setup: lines only save the three sample files, so skip past them.

  1. awk 'NR > 3 ...' adds up us + sy, wa and si + so over the two real samples and counts them in n.
  2. In END, it divides by n to get averages and applies the rules in order.
  3. FILENAME is awk's built-in name of the file it is reading, used in the report line.

cpu.txt has 98% CPU use and no waiting: CPU-bound. io.txt has about 48% I/O wait with nothing swapping: disk-bound. mem.txt has heavy swap traffic, which also explains its I/O wait: memory-bound.

Edge cases. These thresholds are rules of thumb, not laws: a database server may run at 30% wa normally. On a VM, high st means the host is overloaded and the fix is outside your machine. Sample for more than a few seconds, because a spike can mislead.

# Setup: save three vmstat samples in a fresh temporary folder (on a server: vmstat 1 3)
cd "$(mktemp -d)"
cat > cpu.txt << 'OUT'
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 1  0      0 812000  50000 900000    0    0     5    10  200  300  5  2 92  1  0
 8  0      0 810000  50000 900000    0    0     0     4 3000 5000 88 10  2  0  0
 9  0      0 809000  50000 900000    0    0     0     8 3100 5200 90  9  1  0  0
OUT
cat > io.txt << 'OUT'
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 1  0      0 812000  50000 900000    0    0     5    10  200  300  5  2 92  1  0
 1  4      0 790000  50000 910000    0    0  9000 12000  900 1200  6  4 45 45  0
 0  5      0 785000  50000 915000    0    0  9500 11000  950 1300  5  3 42 50  0
OUT
cat > mem.txt << 'OUT'
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
 1  0 400000  90000  2000  60000    5   10    50   80  300  400 10  5 80  5  0
 3  1 520000  30000  1000  40000  300  500   900  1200 1500 2500 20 10 40 30  0
 4  2 600000  25000  1000  35000  450  600  1100  1500 1700 2700 22 12 30 36  0
OUT

for f in cpu.txt io.txt mem.txt; do
  awk 'NR > 3 { cpu += $13 + $14; wa += $16; swap += $7 + $8; n++ }
  END {
    cpu /= n; wa /= n; swap /= n
    if      (swap > 0)  verdict = "memory-bound (swapping)"
    else if (wa >= 20)  verdict = "I/O-bound (waiting on disk)"
    else if (cpu >= 80) verdict = "CPU-bound"
    else                verdict = "healthy"
    printf "%-8s cpu %3.0f%%  wait %3.0f%%  swap %4.0f KiB/s  -> %s\n", FILENAME, cpu, wa, swap, verdict
  }' "$f"
done

Interview follow-ups

  • Sample a live server for one minute and keep only the lines that show a problem.

    Run vmstat 5 13 | awk 'NR > 3 && ($7 + $8 > 0 || $16 >= 20 || $13 + $14 >= 80)'. That takes 13 samples 5 seconds apart, skips the headers and the since-boot line, and prints only the samples that break a rule. Add -t to vmstat (on most versions) to get a timestamp on each line, so you can match problems with log entries or deploys.

Frequently asked questions

Run top and press P to sort by CPU, or ps -eo pid,pcpu,comm --sort=-pcpu | head for a one-shot list. Remember that %CPU in ps is a lifetime average, so top is better for "right now". If sy (kernel time) is high rather than us, the cause is often many small system calls, lots of network traffic, or many processes being created; pidstat 1 or perf top help there.

iotop shows per-process disk reads and writes live, if it is installed. pidstat -d 1 (from the sysstat package) shows the same per process every second. iostat -x 1 shows which disk is busy and its %util and wait times. Without any of them, look for processes in D state with ps -eo stat,pid,cmd | grep '^D'.

Steal time is CPU time the hypervisor gave to other virtual machines while yours wanted to run. Some steal is normal on shared instances; sustained values above about 10% mean your VM is not getting the CPU it pays for. On burstable instance types, high steal can mean you ran out of CPU credits. The fixes are outside the VM: a bigger or dedicated instance type, or a different host.