Bash and Linux

Failed Services from systemd

mediumSystem monitoring

Problem statement

Read saved systemctl --failed and journal output, and for every failed service print the last error message from the program itself. When a server has problems after a reboot or deploy, the first commands are "what failed?" and "why?", and the answer is almost always in the journal.

failed.txt (from systemctl --failed --no-legend --plain)

TEXT
nginx.service loaded failed failed A high performance web server
backup.service loaded failed failed Nightly backup

journal-nginx.service.txt (from journalctl -u nginx.service)

TEXT
Oct 01 09:00:01 web1 systemd[1]: Starting nginx.service - A high performance web server...
Oct 01 09:00:01 web1 nginx[812]: nginx: [emerg] bind() to 0.0.0.0:80 failed (98: Address already in use)
Oct 01 09:00:01 web1 systemd[1]: nginx.service: Control process exited, code=exited, status=1/FAILURE
Oct 01 09:00:01 web1 systemd[1]: nginx.service: Failed with result 'exit-code'.

journal-backup.service.txt (from journalctl -u backup.service)

TEXT
Oct 01 02:30:00 web1 systemd[1]: Starting backup.service - Nightly backup...
Oct 01 02:30:05 web1 backup.sh[901]: rsync: connection refused to backup-host:873
Oct 01 02:30:05 web1 backup.sh[901]: backup failed after 5 seconds
Oct 01 02:30:05 web1 systemd[1]: backup.service: Main process exited, code=exited, status=12/n/a
  1. Print how many services failed.
  2. For each one, print its name, the last line the program itself wrote (not systemd), and the exit status systemd recorded.

Expected output:

TEXT
== failed units: 2 ==
nginx.service
last message: nginx: [emerg] bind() to 0.0.0.0:80 failed (98: Address already in use)
exit: status=1/FAILURE
backup.service
last message: backup failed after 5 seconds
exit: status=12/n/a

Hints

Hint 1: The unit name is field 1 of failed.txt. Read it in a while read loop and open journal-<unit>.txt for each.

Approach

Optimal: systemctl and journal output

Covers: unit states, systemctl --failed, systemctl status, journalctl -u, -b, --since, -p err, reading journal lines, grep -v, grep -o, the usual recovery steps.

Each service is a unit with a state. systemd starts and watches services, which it calls units. Their states:

State Means
active (running) running fine
inactive (dead) stopped, normally on purpose
failed it exited with an error, crashed, or timed out
activating starting up right now

systemctl --failed lists every unit in the failed state. --no-legend --plain drops the header, footer and dots, which makes it easy to parse.

Every journal line has the same shape.

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB L["Oct 01 09:00:01 web1 nginx[812]: nginx: [emerg] bind() failed"]:::gray L --> T["time
Oct 01 09:00:01"]:::blue L --> H["host
web1"]:::gray L --> P["program[PID]
nginx[812]"]:::purple L --> MSG["message
the real reason"]:::red classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px

The part in brackets after the name is the PID. Lines from systemd[1] are systemd describing what it did ("Starting", "Failed with result"); lines from the program itself hold the real reason. That is why the code removes the systemd[1] lines and keeps the program's last words.

The recovery path on a real server.

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB A(["systemctl --failed"]):::purple --> B(["systemctl status name"]):::purple B --> C(["journalctl -u name"]):::purple C --> D["fix the cause"]:::yellow D --> E(["systemctl restart name"]):::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px
Bash
systemctl --failed what failed
systemctl status nginx state, last log lines, exit code
journalctl -u nginx -b --no-pager everything it logged since boot
journalctl -u nginx --since "10 min ago" just the recent part
journalctl -p err -b only errors, from every service
sudo nginx -t check the config (many services have a test mode)
sudo systemctl restart nginx start it again after the fix

Walking through the code. The # Setup: lines only save the sample outputs, so skip past them.

  1. wc -l < failed.txt counts the failed units.
  2. while read -r unit _ reads the first field of each line into unit and the rest into _, a throwaway name.
  3. For each unit, grep -v 'systemd\[1\]:' removes systemd's own lines, tail -n 1 keeps the last program line, and sed 's/.*]: //' cuts everything up to ]: .
  4. grep -o 'status=[^ ,]*' pulls out the exit status from systemd's lines.

nginx failed because port 80 was already in use; backup failed because the backup host refused the connection. Both answers came from the program's own lines.

Edge cases. Some programs log only to their own files, not to the journal; check /var/log/<name>/ if the journal is quiet. The [1] brackets must be escaped as \[1\] in grep, because [1] alone means "the character 1".

# Setup: save sample systemctl and journalctl output in a fresh temporary folder
cd "$(mktemp -d)"
cat > failed.txt << 'OUT'
nginx.service   loaded failed failed A high performance web server
backup.service  loaded failed failed Nightly backup
OUT
cat > journal-nginx.service.txt << 'OUT'
Oct 01 09:00:01 web1 systemd[1]: Starting nginx.service - A high performance web server...
Oct 01 09:00:01 web1 nginx[812]: nginx: [emerg] bind() to 0.0.0.0:80 failed (98: Address already in use)
Oct 01 09:00:01 web1 systemd[1]: nginx.service: Control process exited, code=exited, status=1/FAILURE
Oct 01 09:00:01 web1 systemd[1]: nginx.service: Failed with result 'exit-code'.
OUT
cat > journal-backup.service.txt << 'OUT'
Oct 01 02:30:00 web1 systemd[1]: Starting backup.service - Nightly backup...
Oct 01 02:30:05 web1 backup.sh[901]: rsync: connection refused to backup-host:873
Oct 01 02:30:05 web1 backup.sh[901]: backup failed after 5 seconds
Oct 01 02:30:05 web1 systemd[1]: backup.service: Main process exited, code=exited, status=12/n/a
OUT

echo "== failed units: $(wc -l < failed.txt) =="
while read -r unit _; do
  log="journal-$unit.txt"
  reason=$(grep -v 'systemd\[1\]:' "$log" | tail -n 1 | sed 's/.*]: //')
  status=$(grep -o 'status=[^ ,]*' "$log" | tail -n 1)
  echo "$unit"
  echo "  last message: $reason"
  echo "  exit:         $status"
done < failed.txt

Interview follow-ups

  • Restart each failed service once and report which ones came back.

    Loop over the same unit list: for each one run sudo systemctl restart "$unit", wait a few seconds, then check systemctl is-active --quiet "$unit", which exits 0 only when it is running. Print back or still failing, and for the failing ones print the last journal lines again. Restarting blindly can hide a real problem, so treat this as a first aid step, not a fix, and always read the reason first.

Frequently asked questions

systemd restarts a crashing service if it has Restart= set, but it stops after too many failures in a short time (by default 5 starts in 10 seconds). Then the unit goes to failed and that message appears. The real error is in the lines before it, so read further back with journalctl -u name -n 50. After fixing the cause, run sudo systemctl reset-failed name and start it again.

start and stop change whether it runs now. restart stops and starts it, which drops connections for a moment. reload asks the running program to re-read its config without stopping, if it supports that (nginx does). enable and disable only decide whether it starts at boot; enable --now does both enable and start. A service can be enabled but stopped, or running but not enabled, which surprises people after a reboot.

Filter by priority with -p: journalctl -p err -b shows only error level and worse, from every service, since the last boot. Add -u name for one service and --since "1 hour ago" for a time window. journalctl -f -u name follows new lines live, like tail -f. --no-pager and -o cat (just the message) make output easy to pipe into grep or awk.