Bash and Linux

Run Jobs in the Background

mediumProcess management

Problem statement

Run several jobs at the same time with &, keep each job's PID with $!, wait for each one, and report which ones failed. Running checks, builds or host updates in parallel saves a lot of time, but only if you still notice when one of them fails.

The script defines a fake job that sleeps and then exits with a given code:

job NAME SECONDS EXIT_CODE

TEXT
build sleeps 0.6 s, exits 0
lint sleeps 0.2 s, exits 3
test sleeps 0.4 s, exits 0
  1. Start all three in the background at once, keeping each PID.
  2. Wait for each job in the order they were started, and print whether it passed or failed with its exit code.
  3. Print the order in which they really finished.
  4. Print how many failed.

Expected output:

TEXT
== started 3 jobs, now waiting ==
build: ok
lint: FAILED (exit 3)
test: ok
== the order they really finished in ==
lint finished
test finished
build finished
== summary ==
1 of 3 jobs failed

Hints

Hint 1: command & starts it in the background and $! holds its PID right after. Store it, for example in an array: pids[0]=$!.

Approach

Optimal: & with $! and wait PID

Covers: &, $!, wait, wait PID, exit codes of background jobs, arrays of PIDs, jobs, nohup, disown, xargs -P.

& means "start it and do not wait". Normally the shell waits for each command to finish before running the next. A & at the end starts the command in the background and moves on at once. Three & jobs run side by side, so the total time is about the longest job, not the sum of all three.

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB M(["script"]):::purple --> B["build &
0.6 s"]:::blue M --> L["lint &
0.2 s, exit 3"]:::red M --> T["test &
0.4 s"]:::blue B --> W{{"wait PID
for each one"}}:::yellow L --> W T --> W W --> R["report: 1 of 3 failed"]:::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px

Keep the PID, or you lose the result. Right after starting a job, $! holds its PID. Save it, because $! changes with the next background job. Later, wait PID pauses until that job ends and returns its exit code. That is how you find out it failed.

Command Does
cmd & start cmd in the background
$! PID of the last background job
wait wait for all background jobs; returns 0 (the codes are lost)
wait PID wait for one job; returns that job's exit code
jobs list background jobs of this shell

Plain wait with no PID is the common mistake: it waits for everything but tells you nothing about failures.

Waiting in start order, finishing in any order. The code waits for build, then lint, then test. Waiting for build takes 0.6 seconds, and by then the other two are already done. Their results are kept, so wait returns them straight away. The finish order, recorded in events.log by each job, is lint, test, build: shortest first.

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB subgraph FIN["finish order"] direction TB F1["0.2 s
lint finished"]:::red ~~~ F2["0.4 s
test finished"]:::green ~~~ F3["0.6 s
build finished"]:::green end classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px style FIN fill:transparent,stroke:#6b7280,stroke-width:2px

Walking through the code. The # Setup: lines only define the fake job function, so skip past them.

  1. Each job ... & line starts a job, then saves $! and a name in two arrays at the same index.
  2. The loop runs wait on each PID inside an if. A 0 means ok; otherwise $? holds the job's exit code, which is reported, and failed goes up by one.
  3. events.log shows the real finish order.
  4. The summary prints the number of failures. A real script would end with exit 1 when failed is not 0, so a CI system sees the failure.

Edge cases. Jobs that write to the screen at the same time can mix their lines; send each job's output to its own file instead. If the script exits before waiting, background jobs keep running without anyone checking them. A background job that reads from the keyboard is stopped by the shell.

# Setup: a fake job that sleeps, records when it finished, and exits with a chosen code
cd "$(mktemp -d)"
job() { sleep "$2"; echo "$1 finished" >> events.log; return "$3"; }

job build 0.6 0 & pids[0]=$!; names[0]=build
job lint  0.2 3 & pids[1]=$!; names[1]=lint
job test  0.4 0 & pids[2]=$!; names[2]=test

echo "== started 3 jobs, now waiting =="
failed=0
for i in 0 1 2; do
  if wait "${pids[$i]}"; then
    echo "${names[$i]}: ok"
  else
    code=$?
    echo "${names[$i]}: FAILED (exit $code)"
    failed=$((failed + 1))
  fi
done

echo "== the order they really finished in =="
cat events.log

echo "== summary =="
echo "$failed of 3 jobs failed"

Interview follow-ups

  • Stop all the other jobs as soon as one fails.

    Wait for whichever job ends first with wait -n (bash 4.3+) in a loop. When it returns a non-zero code, send TERM to the PIDs that are still running: kill -TERM "${pids[@]}" 2>/dev/null, then wait for them and exit with the failure code. This "fail fast" pattern saves time in CI, where there is no point finishing the tests when the build already broke.

Frequently asked questions

& runs a job in the background, but it still belongs to your terminal: when you log out, it usually gets a HUP signal and stops. nohup cmd & makes the job ignore HUP and sends its output to nohup.out, so it survives logout. disown removes a job you already started from the shell's job list, with the same effect. For anything long-lived on a server, a systemd service or a tmux session is the better tool.

xargs -P 4 runs up to 4 commands in parallel from a list: printf '%s\n' host1 host2 host3 | xargs -P 4 -I{} ./check.sh {}. xargs exits with 123 if any command failed, so you still notice failures. In bash 4.3+, you can also start jobs in a loop and call wait -n whenever 4 are running, which waits for any one job to finish. GNU parallel does this and more, if it is installed.

Each job writes whenever it is ready, so lines from different jobs interleave in whatever order they happen. Lines can even be split in the middle when output is large. Give each job its own output file, like job > logs/build.log 2>&1 &, and print the files after wait. Prefixing each line with the job name, for example with sed "s/^/[build] /", also helps.