Bash and Linux

Retry with Backoff

mediumBash scripts and automation Must-do

Problem statement

Write a retry function that runs a flaky command until it works, waiting longer after each failure, and gives up after a set number of attempts. Networks, package mirrors and APIs fail for a moment all the time, and a retry with backoff is what makes deploy and CI scripts survive that.

The script uses two fake commands:

commands

TEXT
flaky fails the first 2 times, then succeeds (counts calls in a file)
always_fails never works
  1. Run flaky with retry, allowing 5 attempts and starting with a 1 second wait, doubling each time.
  2. Run always_fails with 2 attempts and show that retry gives up and returns 1.

Expected output:

TEXT
== flaky: up to 5 attempts ==
attempt 1 failed, waiting 1s
attempt 2 failed, waiting 2s
attempt 3 worked
== always_fails: up to 2 attempts ==
attempt 1 failed, waiting 1s
attempt 2 failed, giving up
retry returned 1

Hints

Hint 1: until "$@"; do ...; done keeps running the command passed to the function until it succeeds. Inside the loop, sleep "$delay" and then delay=$((delay * 2)).

Approach

Optimal: until loop with doubling delay

Covers: functions, local, arguments "$@" in functions, until loops, sleep, $(( )), return codes, exponential backoff, jitter, which errors are worth retrying.

Strict mode, used on every page in this section. The second line, set -euo pipefail, makes bash stop on mistakes instead of carrying on:

Option Means
-e exit as soon as a command fails (with some exceptions, see Handle Command Failures)
-u treat an unset variable as an error, instead of silently using empty text
-o pipefail a pipeline fails if any command in it fails, not just the last one

Put it right after the shebang line #!/usr/bin/env bash in every script you write.

Functions work like small scripts. retry 5 1 flaky calls the function with three arguments: $1 is the attempt limit, $2 the first delay, and the rest is the command. After shift 2, "$@" is exactly the command and its own arguments, so retry can run anything. local keeps the function's variables from leaking into the rest of the script.

until runs until something succeeds. until "$@"; do ...; done runs the command, and if it fails, runs the loop body and tries again. The command's exit code is the condition, so set -e does not stop the script when it fails.

Backoff: wait longer each time. Retrying at once usually hits the same problem. Waiting 1, then 2, then 4 seconds gives the other side time to recover, and stops hundreds of clients from hammering a struggling server together:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB A1["attempt 1 fails"]:::red --> W1(["wait 1 s"]):::yellow W1 --> A2["attempt 2 fails"]:::red A2 --> W2(["wait 2 s"]):::yellow W2 --> A3["attempt 3 works"]:::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px

Always cap it. A retry without a limit can hang a deploy forever. The function counts attempts and returns 1 when it runs out, so the caller can stop or report.

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB R(["run the command"]):::purple --> OK{{"worked?"}}:::yellow OK --> D["return 0"]:::green OK --> L{{"attempts left?"}}:::yellow L --> W(["sleep, double the delay"]):::blue W --> R L --> G["return 1: give up"]:::red classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px

Walking through the code. The # Setup: lines only define the fake commands, so skip past them. flaky keeps its call count in a file, because it is run fresh each time.

  1. retry 5 1 flaky: attempt 1 fails, wait 1 s; attempt 2 fails, wait 2 s; attempt 3 works. The function prints each step.
  2. retry 2 1 always_fails: attempt 1 fails, wait 1 s; attempt 2 fails, and the limit is reached. || echo shows the return code 1.

The page waits about 4 seconds in total, because the sleeps are real.

Edge cases. Only retry errors that can go away: a timeout or HTTP 503 may pass, but a 404, a wrong password or a syntax error never will, so retrying just wastes time. Commands that change things should be safe to repeat; retrying a "create user" that half-worked can create two.

#!/usr/bin/env bash
set -euo pipefail

# Setup: two fake commands in a fresh temporary folder
cd "$(mktemp -d)"
echo 0 > calls
flaky() {                         # fails twice, then works
  local n; n=$(( $(cat calls) + 1 )); echo "$n" > calls
  (( n >= 3 ))
}
always_fails() { return 7; }

# retry MAX_ATTEMPTS FIRST_DELAY COMMAND [ARGS...]
retry() {
  local max=$1 delay=$2 attempt=1
  shift 2
  until "$@"; do
    if (( attempt >= max )); then
      echo "  attempt $attempt failed, giving up"
      return 1
    fi
    echo "  attempt $attempt failed, waiting ${delay}s"
    sleep "$delay"
    attempt=$((attempt + 1))
    delay=$((delay * 2))
  done
  echo "  attempt $attempt worked"
}

echo "== flaky: up to 5 attempts =="
retry 5 1 flaky

echo "== always_fails: up to 2 attempts =="
retry 2 1 always_fails || echo "retry returned $?"
RecapThe whole problem in a few lines, for the night before
  • Spot it: "flaky command", "try again a few times"
  • Idea: until "$@"; do sleep $delay; delay=$((delay * 2)); done with an attempt cap
  • Cost: the total wait grows quickly: 1 + 2 + 4 + 8 seconds
  • Trap: retrying forever, or retrying errors that will never pass (404, bad password)

Interview follow-ups

  • Give up after a total time limit instead of an attempt count.

    Record the start with start=$SECONDS (bash's built-in seconds counter), and in the loop check (( SECONDS - start + delay > limit )) before sleeping, returning 1 if the next wait would pass the limit. You can keep the attempt limit too and stop at whichever comes first. For a hard limit on a single attempt, wrap the command itself in timeout 10, so one hung attempt cannot use all the time.

Frequently asked questions

If a thousand servers all fail at the same moment and all retry after exactly 1, 2 and 4 seconds, they hit the recovering service together each time. Jitter adds a small random amount to each wait, spreading the retries out. In bash, sleep "$(( delay + RANDOM % delay ))" waits between delay and twice that. Big cloud providers recommend backoff with jitter for exactly this reason.

Often not. curl --retry 5 --retry-delay 1 retries network errors with backoff built in, and --retry-all-errors widens it. wget --tries, package managers and many CLIs have similar options. Kubernetes, systemd (Restart=on-failure with RestartSec=) and CI systems retry at their own level. Write your own retry when a step has no built-in option, like a custom health check or a git push in CI.

It depends on how long the problem usually lasts and how long the caller can wait. For a network blip, 3 to 5 attempts starting at 1 second, doubling, covers about half a minute. Also set a maximum single wait, like 30 seconds, so the doubling does not grow to hours. For a deploy step, think about the total time: 5 attempts with doubling from 1 second can take over 30 seconds before it gives up.