Failed Services from systemd
Problem statement
Read saved systemctl --failed and journal output, and for every failed service print the last error message from the program itself. When a server has problems after a reboot or deploy, the first commands are "what failed?" and "why?", and the answer is almost always in the journal.
failed.txt (from systemctl --failed --no-legend --plain)
nginx.service loaded failed failed A high performance web serverbackup.service loaded failed failed Nightly backupjournal-nginx.service.txt (from journalctl -u nginx.service)
Oct 01 09:00:01 web1 systemd[1]: Starting nginx.service - A high performance web server...Oct 01 09:00:01 web1 nginx[812]: nginx: [emerg] bind() to 0.0.0.0:80 failed (98: Address already in use)Oct 01 09:00:01 web1 systemd[1]: nginx.service: Control process exited, code=exited, status=1/FAILUREOct 01 09:00:01 web1 systemd[1]: nginx.service: Failed with result 'exit-code'.journal-backup.service.txt (from journalctl -u backup.service)
Oct 01 02:30:00 web1 systemd[1]: Starting backup.service - Nightly backup...Oct 01 02:30:05 web1 backup.sh[901]: rsync: connection refused to backup-host:873Oct 01 02:30:05 web1 backup.sh[901]: backup failed after 5 secondsOct 01 02:30:05 web1 systemd[1]: backup.service: Main process exited, code=exited, status=12/n/a- Print how many services failed.
- For each one, print its name, the last line the program itself wrote (not systemd), and the exit status systemd recorded.
Expected output:
== failed units: 2 ==nginx.service last message: nginx: [emerg] bind() to 0.0.0.0:80 failed (98: Address already in use) exit: status=1/FAILUREbackup.service last message: backup failed after 5 seconds exit: status=12/n/aHints
failed.txt. Read it in a while read loop and open journal-<unit>.txt for each.Approach
Optimal: systemctl and journal output
Covers: unit states, systemctl --failed, systemctl status, journalctl -u, -b, --since, -p err, reading journal lines, grep -v, grep -o, the usual recovery steps.
Each service is a unit with a state. systemd starts and watches services, which it calls units. Their states:
| State | Means |
|---|---|
active (running) |
running fine |
inactive (dead) |
stopped, normally on purpose |
failed |
it exited with an error, crashed, or timed out |
activating |
starting up right now |
systemctl --failed lists every unit in the failed state. --no-legend --plain drops the header, footer and dots, which makes it easy to parse.
Every journal line has the same shape.
Oct 01 09:00:01"]:::blue L --> H["host
web1"]:::gray L --> P["program[PID]
nginx[812]"]:::purple L --> MSG["message
the real reason"]:::red classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px
The part in brackets after the name is the PID. Lines from systemd[1] are systemd describing what it did ("Starting", "Failed with result"); lines from the program itself hold the real reason. That is why the code removes the systemd[1] lines and keeps the program's last words.
The recovery path on a real server.
systemctl --failed what failedsystemctl status nginx state, last log lines, exit codejournalctl -u nginx -b --no-pager everything it logged since bootjournalctl -u nginx --since "10 min ago" just the recent partjournalctl -p err -b only errors, from every servicesudo nginx -t check the config (many services have a test mode)sudo systemctl restart nginx start it again after the fixWalking through the code. The # Setup: lines only save the sample outputs, so skip past them.
wc -l < failed.txtcounts the failed units.while read -r unit _reads the first field of each line intounitand the rest into_, a throwaway name.- For each unit,
grep -v 'systemd\[1\]:'removes systemd's own lines,tail -n 1keeps the last program line, andsed 's/.*]: //'cuts everything up to]:. grep -o 'status=[^ ,]*'pulls out the exit status from systemd's lines.
nginx failed because port 80 was already in use; backup failed because the backup host refused the connection. Both answers came from the program's own lines.
Edge cases. Some programs log only to their own files, not to the journal; check /var/log/<name>/ if the journal is quiet. The [1] brackets must be escaped as \[1\] in grep, because [1] alone means "the character 1".
# Setup: save sample systemctl and journalctl output in a fresh temporary folder
cd "$(mktemp -d)"
cat > failed.txt << 'OUT'
nginx.service loaded failed failed A high performance web server
backup.service loaded failed failed Nightly backup
OUT
cat > journal-nginx.service.txt << 'OUT'
Oct 01 09:00:01 web1 systemd[1]: Starting nginx.service - A high performance web server...
Oct 01 09:00:01 web1 nginx[812]: nginx: [emerg] bind() to 0.0.0.0:80 failed (98: Address already in use)
Oct 01 09:00:01 web1 systemd[1]: nginx.service: Control process exited, code=exited, status=1/FAILURE
Oct 01 09:00:01 web1 systemd[1]: nginx.service: Failed with result 'exit-code'.
OUT
cat > journal-backup.service.txt << 'OUT'
Oct 01 02:30:00 web1 systemd[1]: Starting backup.service - Nightly backup...
Oct 01 02:30:05 web1 backup.sh[901]: rsync: connection refused to backup-host:873
Oct 01 02:30:05 web1 backup.sh[901]: backup failed after 5 seconds
Oct 01 02:30:05 web1 systemd[1]: backup.service: Main process exited, code=exited, status=12/n/a
OUT
echo "== failed units: $(wc -l < failed.txt) =="
while read -r unit _; do
log="journal-$unit.txt"
reason=$(grep -v 'systemd\[1\]:' "$log" | tail -n 1 | sed 's/.*]: //')
status=$(grep -o 'status=[^ ,]*' "$log" | tail -n 1)
echo "$unit"
echo " last message: $reason"
echo " exit: $status"
done < failed.txtInterview follow-ups
Restart each failed service once and report which ones came back.
Loop over the same unit list: for each one run
sudo systemctl restart "$unit", wait a few seconds, then checksystemctl is-active --quiet "$unit", which exits 0 only when it is running. Printbackorstill failing, and for the failing ones print the last journal lines again. Restarting blindly can hide a real problem, so treat this as a first aid step, not a fix, and always read the reason first.
Frequently asked questions
systemd restarts a crashing service if it has Restart= set, but it stops after too many failures in a short time (by default 5 starts in 10 seconds). Then the unit goes to failed and that message appears. The real error is in the lines before it, so read further back with journalctl -u name -n 50. After fixing the cause, run sudo systemctl reset-failed name and start it again.
start and stop change whether it runs now. restart stops and starts it, which drops connections for a moment. reload asks the running program to re-read its config without stopping, if it supports that (nginx does). enable and disable only decide whether it starts at boot; enable --now does both enable and start. A service can be enabled but stopped, or running but not enabled, which surprises people after a reboot.
Filter by priority with -p: journalctl -p err -b shows only error level and worse, from every service, since the last boot. Add -u name for one service and --since "1 hour ago" for a time window. journalctl -f -u name follows new lines live, like tail -f. --no-pager and -o cat (just the message) make output easy to pipe into grep or awk.