Bash and Linux

Find and Stop a Process

mediumProcess management Must-do

Problem statement

Find a running process by name, stop it politely with TERM, and if it ignores that, force it with KILL. This is a standard task when a runaway job fills a disk or a stuck service must be restarted, and doing it in the right order lets programs clean up instead of leaving broken files behind.

The script makes two small worker programs and runs each in the background:

demo-worker (cleans up and exits when it gets TERM)

Bash
#!/bin/bash
trap 'echo "worker: got TERM, cleaning up"; exit 0' TERM
while :; do sleep 0.1; done

stubborn-worker (ignores TERM)

TEXT
#!/bin/bash
trap '' TERM
while :; do sleep 0.1; done
  1. Find demo-worker by its exact name and count how many are running.
  2. Stop it with TERM, wait for it, and print its exit code and the new count.
  3. Do the same with stubborn-worker: TERM is ignored, so after about a second send KILL.

Expected output:

TEXT
== find it by name ==
demo-worker processes: 1
== stop it politely with TERM ==
worker: got TERM, cleaning up
exit code: 0
demo-worker processes: 0
== a worker that ignores TERM ==
still running after TERM, sending KILL
exit code: 137 (128 + 9, killed by KILL)
stubborn-worker processes: 0

Hints

Hint 1: pgrep -x name prints the PIDs of processes whose name is exactly name. kill -0 PID sends no signal; it only checks whether the process still exists.

Approach

Optimal: TERM, wait, then KILL

Covers: signals, kill -TERM, kill -KILL, kill -0, pgrep -x, pgrep -f, pkill, trap, $!, wait, exit codes above 128, the [n]ame grep trick.

A signal is a message to a process. kill does not only kill; it sends a signal, and the process decides what to do with most of them:

Signal Number Means Can the process catch it?
TERM 15 please stop (the default for kill) yes: it can clean up first
INT 2 interrupt, what Ctrl+C sends yes
HUP 1 hang up; many daemons reload their config yes
KILL 9 stop now no: the kernel removes it
STOP / CONT 19 / 18 pause and resume no / yes

Always try TERM first. A program that gets TERM can finish writing, close database connections and delete its lock files. KILL gives it no chance, so half-written files and stale locks are common after a kill -9. The safe order:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB F(["pgrep -x name"]):::blue --> T(["kill -TERM PID"]):::purple T --> W{{"gone within a second?"}}:::yellow W --> OK["done: it cleaned up"]:::green W --> K(["kill -KILL PID"]):::red K --> G["gone, no cleanup"]:::red classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px

Finding the PID. pgrep -x demo-worker matches the process name exactly. pgrep -f text matches anywhere in the full command line, which is handy but can match more than you meant, including your own shell. pkill takes the same options and sends the signal directly.

The grep trick. Many people find processes with ps aux | grep nginx. The grep nginx process itself is also running and contains the word, so it shows up in its own results. Writing the pattern as [n]ginx fixes it: the pattern still matches the text nginx, but the grep's own command line now contains [n]ginx, which does not match. pgrep avoids the problem completely.

TEXT
ps aux | grep nginx shows nginx, plus "grep nginx" itself
ps aux | grep '[n]ginx' shows only nginx
pgrep -x nginx prints only the PIDs

Exit codes tell you how it ended. After wait PID, $? is the process's exit code. 0 means it exited cleanly after catching TERM. A number above 128 means it was killed by a signal: 128 plus the signal number, so 137 is KILL (9) and 143 is TERM (15) when not caught.

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart LR subgraph CODES["wait PID, then $?"] direction TB C0["0
exited cleanly"]:::green ~~~ C143["143 = 128 + 15
TERM, not caught"]:::yellow ~~~ C137["137 = 128 + 9
KILL"]:::red end classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px style CODES fill:transparent,stroke:#6b7280,stroke-width:2px

Walking through the code. The # Setup: lines only write the two worker scripts and make them runnable, so skip past them.

  1. ./demo-worker & starts it in the background, and sleep 0.3 gives it time to start. pgrep -x demo-worker | wc -l counts 1.
  2. stop_gracefully sends TERM, then checks up to 10 times with kill -0, a tenth of a second apart. The worker's trap prints its message and exits 0, so the function returns early.
  3. stubborn-worker ignores TERM, so after the checks the function sends KILL. wait reports 137.

The { ... } 2>/dev/null around the last part hides bash's own "Killed" notice, which it writes to stderr when one of its background jobs is killed.

Edge cases. If pgrep finds nothing, $pid is empty and kill prints a usage error; check with [[ -n $pid ]] first. If it finds several, kill them all with pkill -TERM -x name. You can only signal your own processes unless you are root.

# Setup: two small worker programs in a fresh temporary folder
cd "$(mktemp -d)"
cat > demo-worker << 'W'
#!/bin/bash
trap 'echo "worker: got TERM, cleaning up"; exit 0' TERM
while :; do sleep 0.1; done
W
cat > stubborn-worker << 'W'
#!/bin/bash
trap '' TERM
while :; do sleep 0.1; done
W
chmod +x demo-worker stubborn-worker

# Ask politely with TERM; if the process is still there after about 1 second, send KILL
stop_gracefully() {
  local pid=$1
  kill -TERM "$pid"
  for _ in 1 2 3 4 5 6 7 8 9 10; do
    kill -0 "$pid" 2>/dev/null || return 0   # kill -0 only checks that it exists
    sleep 0.1
  done
  echo "still running after TERM, sending KILL"
  kill -KILL "$pid"
}

./demo-worker &
sleep 0.3
echo "== find it by name =="
echo "demo-worker processes: $(pgrep -x demo-worker | wc -l)"
pid=$(pgrep -x demo-worker)

echo "== stop it politely with TERM =="
stop_gracefully "$pid"
wait "$pid" 2>/dev/null; echo "exit code: $?"
echo "demo-worker processes: $(pgrep -x demo-worker | wc -l)"

# The { } 2>/dev/null hides bash's own "Killed" notice, which it prints to stderr
{
  ./stubborn-worker &
  sleep 0.3
  pid=$(pgrep -x stubborn-worker)
  echo "== a worker that ignores TERM =="
  stop_gracefully "$pid"
  wait "$pid"; echo "exit code: $? (128 + 9, killed by KILL)"
  echo "stubborn-worker processes: $(pgrep -x stubborn-worker | wc -l)"
} 2>/dev/null
RecapThe whole problem in a few lines, for the night before
  • Spot it: "stop the runaway process", "it will not die"
  • Idea: pgrep -x name, kill -TERM, check with kill -0, KILL only if still alive
  • Cost: a second or so of waiting gives the program time to clean up
  • Trap: kill -9 first (no cleanup), or ps | grep matching the grep itself

Interview follow-ups

  • Stop a whole group of processes, like a script and every child it started.

    Children usually share the parent's process group, so you can signal the whole group by passing a negative group ID: kill -TERM -- -<PGID>. Find the group with ps -o pgid= -p <PID>. pkill -TERM -P <PID> signals only the direct children. Inside your own scripts, start workers with setsid if you want them in a separate group you can stop together.

Frequently asked questions

Use sudo systemctl stop name instead of kill. systemd sends TERM to every process of the service, waits (90 seconds by default, set by TimeoutStopSec=), and then sends KILL, which is the same pattern as this page. It also knows the service is stopped on purpose, so it does not restart it. Killing a service's process by hand often just makes systemd start it again if it has Restart= set.

You can only send signals to your own processes. Another user's process, or one run by root, needs sudo kill. If it still fails as root, the process may be in D state, waiting on a stuck disk or network mount, and even KILL waits until that call returns. Zombies (Z) also ignore signals, because they have already finished; fix their parent instead.

lsof /var/log/bad.log lists every process that has the file open, with its PID and command. fuser -v /var/log/bad.log does the same in a shorter form, and fuser -k can send them a signal. On a system without either, look through /proc/*/fd for links that point at the file. Once you have the PID, stop it with TERM first, as on this page.