Bash and Linux

Flag Full Disks

easyDisk and storage Must-do

Problem statement

Read df -P output and list the real disks that are more than 85% full, fullest first, with the free space left. "Disk almost full" is one of the most common alerts an on-call engineer gets, and this is exactly the check behind it.

df.txt (from df -P; sizes are in 1024-byte blocks)

Bash
Filesystem 1024-blocks Used Available Capacity Mounted on
/dev/nvme0n1p1 50000000 46000000 4000000 92% /
tmpfs 4000000 3900000 100000 98% /dev/shm
/dev/nvme1n1 200000000 176000000 24000000 88% /data
/dev/mapper/vg0-logs 20000000 9000000 11000000 45% /var/log
overlay 50000000 46000000 4000000 92% /var/lib/docker/overlay2/abc/merged

Skip memory and container filesystems (tmpfs, devtmpfs, overlay). For the rest, print every mount above 85%, fullest first, as use% mount free GiB.

Expected output:

TEXT
== disks above 85%, fullest first ==
92% / 3.8 GiB free
88% /data 22.9 GiB free

Hints

Hint 1: Skip the header with NR > 1 and skip lines whose field 1 is tmpfs, devtmpfs or overlay. The use % is field 5, the free space field 4 and the mount point field 6.

Approach

Optimal: df -P with awk

Covers: df -h, df -P, df columns, skipping pseudo filesystems, turning 92% into a number, sort -rn, printf, reserved blocks.

What df reports. df (disk free) lists every mounted filesystem with its size, how much is used and how much is left:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB L["/dev/nvme0n1p1 50000000 46000000 4000000 92% /"]:::gray L --> F1["$1 device"]:::blue L --> F4["$4 available KiB"]:::green L --> F5["$5 use %"]:::red L --> F6["$6 mount"]:::purple classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px
Column Field Means
Filesystem 1 the device or kind: /dev/nvme0n1p1, tmpfs, overlay
1024-blocks 2 size in KiB
Used 3 used, KiB
Available 4 free for normal users, KiB
Capacity 5 use %, like 92%
Mounted on 6 the folder where it appears

df -h prints sizes like 46G, which is nicer to read but harder to compute with. In scripts, use plain KiB numbers.

Why -P? Without -P, df may split a line in two when the device name is long, like /dev/mapper/vg0-logs: the name on one line and the numbers on the next. Every field number after that is wrong. -P (POSIX format) promises one line per filesystem.

Skip what is not a disk. tmpfs and devtmpfs live in RAM, and overlay is a container's view of a real disk, so it repeats the root disk's numbers. Alerting on them causes noise or double alerts. In the sample, the overlay line shows the same 92% as /, and /dev/shm shows 98% but is only memory.

Text vs number. Field 5 is the text 92%. Compared as text, "9%" is bigger than "85", because 9 comes after 8. $5 + 0 makes awk read the number at the start, 92, and drop the %. Always force numbers before comparing.

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart LR subgraph TXT["as text"] direction TB T1[""9%" vs "85""]:::red --> T2["9 is after 8
so 9% looks bigger"]:::red end subgraph NUM["as a number: $5 + 0"] direction TB N1["9 vs 85"]:::green --> N2["9 is smaller
correct"]:::green end TXT ~~~ NUM classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px style TXT fill:transparent,stroke:#dc2626,stroke-width:2px style NUM fill:transparent,stroke:#059669,stroke-width:2px

Walking through the code. The # Setup: lines only save the sample output, so skip past them. On a server you would pipe df -P straight into the awk.

  1. NR > 1 skips the header, and $1 !~ /^(tmpfs|devtmpfs|overlay)$/ skips pseudo filesystems.
  2. $5 + 0 > 85 keeps the full ones, and printf prints the use %, mount and free GiB.
  3. sort -rn puts the fullest first.

Edge cases. ext4 keeps about 5% of a disk reserved for root by default, so a disk can show 95% and already be full for normal users; Available already accounts for that. A disk can also be "full" with free space when it runs out of inodes; that is the next page.

# Setup: save sample df -P output in a fresh temporary folder (on a server: df -P | awk ...)
cd "$(mktemp -d)"
cat > df.txt << 'OUT'
Filesystem                1024-blocks      Used Available Capacity Mounted on
/dev/nvme0n1p1               50000000  46000000   4000000      92% /
tmpfs                         4000000   3900000    100000      98% /dev/shm
/dev/nvme1n1                200000000 176000000  24000000      88% /data
/dev/mapper/vg0-logs         20000000   9000000  11000000      45% /var/log
overlay                      50000000  46000000   4000000      92% /var/lib/docker/overlay2/abc/merged
OUT

echo "== disks above 85%, fullest first =="
awk 'NR > 1 && $1 !~ /^(tmpfs|devtmpfs|overlay)$/ && $5 + 0 > 85 {
       printf "%d%%  %-6s %.1f GiB free\n", $5, $6, $4 / 1048576
     }' df.txt | sort -rn
RecapThe whole problem in a few lines, for the night before
  • Spot it: "which disks are almost full"
  • Idea: df -P | awk 'NR > 1 && $5 + 0 > 85', skipping tmpfs and overlay
  • Cost: instant: df reads filesystem totals, not files
  • Trap: comparing 92% as text, or parsing df without -P

Interview follow-ups

  • Turn this into a cron check that exits 1 and prints the mounts when any disk is above the limit.

    Capture the output: full=$(df -P | awk 'NR > 1 && $1 !~ /^(tmpfs|devtmpfs|overlay)$/ && $5 + 0 > 85 {print $5, $6}'). Then if [ -n "$full" ]; then echo "$full"; exit 1; fi. Cron mails any output to the job's owner when mail is set up, and monitoring agents treat exit code 1 as a warning. Pass the limit in as a variable with awk -v limit=85, so the same script works for different servers.

Frequently asked questions

Find what grew: sudo du -xh --max-depth=1 / | sort -rh | head shows the biggest top-level folders on that disk, and the Biggest Directories page drills down. The usual suspects are logs in /var/log, old Docker images and containers (docker system df), package caches, core dumps and temp files. If du finds much less than df reports, look for deleted files still held open, covered two pages on.

The console shows the size of the block device you bought, while df shows the filesystem on it. They differ when the filesystem was not grown after the volume was resized (growpart and resize2fs or xfs_growfs fix that), because of filesystem overhead and reserved blocks, and because consoles often use GB (1000) while df -h uses GiB (1024). Check lsblk for the device size and df -h for the filesystem size.

A fixed threshold like 85% is a start, but a large disk at 85% may have weeks of room and a small one only minutes. Better alerts look at the rate: how fast is it growing, and when will it hit 100%? Monitoring systems do this with "predict linear" style rules. In a script, you can store yesterday's Used value and compare. Alert on inodes too, which this check does not see.