Bash and Linux

Check Memory and Load

mediumSystem monitoring Must-do

Problem statement

Read saved free -m and /proc/loadavg output and decide whether the machine is short of memory or overloaded, the way a health check or an on-call engineer would. Both numbers are easy to misread: "free" memory is almost always low on Linux, and a load of 6 can be fine or terrible depending on the CPU count.

free.txt (from free -m)

TEXT
total used free shared buff/cache available
Mem: 7820 5210 310 120 2300 2350
Swap: 2047 900 1147

loadavg.txt (from cat /proc/loadavg) on a machine with 4 CPUs

TEXT
6.20 4.10 2.50 3/412 98765
  1. Print memory used % the wrong way (from free) and the right way (from available).
  2. Print swap use and warn if any swap is in use.
  3. Print the three load averages, the load per CPU, and warn if the 1-minute load is above the CPU count.

Expected output:

TEXT
== memory ==
used (wrong, from free): 96.0%
used (right, from available): 69.9%
swap used: 900 of 2047 MiB
WARN: swap is in use
== load ==
load 1/5/15 min: 6.20 4.10 2.50 on 4 CPUs
per CPU now: 1.55
WARN: more tasks than CPUs, and rising since 15 min ago

Hints

Hint 1: On Linux, unused RAM is used as file cache, so free is always small. The available column is how much RAM programs can still get. Used % is (total - available) / total.

Approach

Optimal: free and loadavg with awk

Covers: free -m columns, available vs free, swap, /proc/loadavg, nproc, awk maths and printf, comparing decimals in awk (not in [ ]).

Linux uses spare RAM as a cache. When memory is not needed by programs, Linux fills it with copies of files it has read, so the next read is fast. That cache is handed back the moment a program needs memory. So the free column looks tiny on a healthy machine:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart LR subgraph RAM["7820 MiB of RAM"] direction TB U["used by programs
5210"]:::red ~~~ C["file cache
2300, can be given back"]:::yellow ~~~ F["free
310"]:::gray end RAM --> AV["available ≈ free + most cache
2350"]:::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px style RAM fill:transparent,stroke:#6b7280,stroke-width:2px
Column Means
total all RAM
used used by programs
free not used for anything at all, usually small
buff/cache file cache, can be given back
available what programs can still get: free plus most of the cache

Here free says 96% used, which sounds like an emergency. available says 69.9%, which is fine. Alerts should always use available.

Swap. Swap is disk space used as overflow RAM. A little swap used in the past is not a problem. Swap being read and written right now (si/so in vmstat, next page) is, because disk is thousands of times slower than RAM.

Load average. /proc/loadavg starts with three numbers: the average number of tasks running or waiting to run over 1, 5 and 15 minutes. On Linux it also counts tasks waiting for disk. Compare it with the number of CPUs:

Load per CPU Means
below 0.7 plenty of room
about 1.0 busy, every CPU working
above 1.0 tasks are waiting their turn

6.20 4.10 2.50 on 4 CPUs is 1.55 per CPU now, and it was lower 5 and 15 minutes ago, so load is rising. That trend is the key thing to report.

Comparing decimals. Bash's [ ] and (( )) only handle whole numbers, so [ 6.20 -gt 4 ] fails. awk understands decimals, which is why the checks are done in awk.

Walking through the code. The # Setup: lines only save the sample outputs and the CPU count, so skip past them. On a real server you would run free -m, cat /proc/loadavg and nproc instead.

  1. The first awk picks the Mem: line, prints both percentages, then the Swap: line and a warning when field 3 is above 0.
  2. The second awk takes the CPU count with -v, prints the three loads and the per-CPU value, and warns when the 1-minute load is higher than the CPUs.
%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB L["load 6.20"]:::red --> D(["divide by 4 CPUs"]):::purple D --> P["1.55 per CPU"]:::yellow P --> W["above 1.0: tasks are waiting"]:::red classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px

Edge cases. Old versions of free have no available column; on those, read MemAvailable from /proc/meminfo. In containers, free shows the host's memory, not the container's limit; check /sys/fs/cgroup/memory.max instead.

# Setup: save sample outputs in a fresh temporary folder (on a server: free -m, /proc/loadavg, nproc)
cd "$(mktemp -d)"
cat > free.txt << 'OUT'
               total        used        free      shared  buff/cache   available
Mem:            7820        5210         310         120        2300        2350
Swap:           2047         900        1147
OUT
echo '6.20 4.10 2.50 3/412 98765' > loadavg.txt
cpus=4

echo "== memory =="
awk '
  $1 == "Mem:"  { printf "used (wrong, from free):      %.1f%%\n", ($2 - $4) / $2 * 100
                  printf "used (right, from available): %.1f%%\n", ($2 - $7) / $2 * 100 }
  $1 == "Swap:" { printf "swap used: %d of %d MiB\n", $3, $2
                  if ($3 > 0) print "WARN: swap is in use" }
' free.txt

echo "== load =="
awk -v cpus="$cpus" '{
  printf "load 1/5/15 min: %s %s %s on %d CPUs\n", $1, $2, $3, cpus
  printf "per CPU now: %.2f\n", $1 / cpus
  if ($1 > cpus) print "WARN: more tasks than CPUs, and rising since 15 min ago"
}' loadavg.txt
RecapThe whole problem in a few lines, for the night before
  • Spot it: "is the box out of memory", "is it overloaded"
  • Idea: memory used % from available, not free; load average compared with nproc
  • Cost: reading two small files
  • Trap: alerting on the free column, or comparing decimals with [ ]

Interview follow-ups

  • Turn this into a check that exits 1 when either warning fires, for cron or a monitoring agent.

    Have each awk print a word like WARN and count them: warnings=$( { awk ... free.txt; awk ... loadavg.txt; } | grep -c '^WARN' ). Then [ "$warnings" -eq 0 ] || exit 1. Monitoring tools like Nagios-style checks use exit codes 0 for OK, 1 for warning and 2 for critical, so returning the right code lets the same script plug into them. Read live values from /proc instead of saved files.

Frequently asked questions

On Linux, the load average also counts tasks in D state, waiting for disk or network storage. A slow disk or a hung NFS mount can push load to 50 while the CPUs sit idle. Check vmstat 1 for a high wa (I/O wait) and b (blocked) column, and ps -eo stat,pid,cmd | grep '^D' for the stuck processes. The fix is in storage, not in adding CPUs.

free -m (or free -h for human units) shows memory, uptime shows the three load averages with the time, and nproc prints the CPU count. cat /proc/meminfo and cat /proc/loadavg are the raw sources these tools read. top shows them all at once in its header and updates every few seconds. In a script, reading /proc directly avoids depending on the output format of tools.

No. Linux may move memory that was used once, at startup for example, into swap and leave it there, which frees RAM for cache. That shows as "swap used" but costs nothing. The danger is active swapping, when memory moves in and out all the time and everything slows down. Watch si and so in vmstat 1: steady non-zero values mean the machine needs more RAM or a process is using too much.