Retry with Backoff
Problem statement
Write a retry function that runs a flaky command until it works, waiting longer after each failure, and gives up after a set number of attempts. Networks, package mirrors and APIs fail for a moment all the time, and a retry with backoff is what makes deploy and CI scripts survive that.
The script uses two fake commands:
commands
flaky fails the first 2 times, then succeeds (counts calls in a file)always_fails never works- Run
flakywithretry, allowing 5 attempts and starting with a 1 second wait, doubling each time. - Run
always_failswith 2 attempts and show thatretrygives up and returns 1.
Expected output:
== flaky: up to 5 attempts == attempt 1 failed, waiting 1s attempt 2 failed, waiting 2s attempt 3 worked== always_fails: up to 2 attempts == attempt 1 failed, waiting 1s attempt 2 failed, giving upretry returned 1Hints
until "$@"; do ...; done keeps running the command passed to the function until it succeeds. Inside the loop, sleep "$delay" and then delay=$((delay * 2)).Approach
Optimal: until loop with doubling delay
Covers: functions, local, arguments "$@" in functions, until loops, sleep, $(( )), return codes, exponential backoff, jitter, which errors are worth retrying.
Strict mode, used on every page in this section. The second line, set -euo pipefail, makes bash stop on mistakes instead of carrying on:
| Option | Means |
|---|---|
-e |
exit as soon as a command fails (with some exceptions, see Handle Command Failures) |
-u |
treat an unset variable as an error, instead of silently using empty text |
-o pipefail |
a pipeline fails if any command in it fails, not just the last one |
Put it right after the shebang line #!/usr/bin/env bash in every script you write.
Functions work like small scripts. retry 5 1 flaky calls the function with three arguments: $1 is the attempt limit, $2 the first delay, and the rest is the command. After shift 2, "$@" is exactly the command and its own arguments, so retry can run anything. local keeps the function's variables from leaking into the rest of the script.
until runs until something succeeds. until "$@"; do ...; done runs the command, and if it fails, runs the loop body and tries again. The command's exit code is the condition, so set -e does not stop the script when it fails.
Backoff: wait longer each time. Retrying at once usually hits the same problem. Waiting 1, then 2, then 4 seconds gives the other side time to recover, and stops hundreds of clients from hammering a struggling server together:
Always cap it. A retry without a limit can hang a deploy forever. The function counts attempts and returns 1 when it runs out, so the caller can stop or report.
Walking through the code. The # Setup: lines only define the fake commands, so skip past them. flaky keeps its call count in a file, because it is run fresh each time.
retry 5 1 flaky: attempt 1 fails, wait 1 s; attempt 2 fails, wait 2 s; attempt 3 works. The function prints each step.retry 2 1 always_fails: attempt 1 fails, wait 1 s; attempt 2 fails, and the limit is reached.|| echoshows the return code 1.
The page waits about 4 seconds in total, because the sleeps are real.
Edge cases. Only retry errors that can go away: a timeout or HTTP 503 may pass, but a 404, a wrong password or a syntax error never will, so retrying just wastes time. Commands that change things should be safe to repeat; retrying a "create user" that half-worked can create two.
#!/usr/bin/env bash
set -euo pipefail
# Setup: two fake commands in a fresh temporary folder
cd "$(mktemp -d)"
echo 0 > calls
flaky() { # fails twice, then works
local n; n=$(( $(cat calls) + 1 )); echo "$n" > calls
(( n >= 3 ))
}
always_fails() { return 7; }
# retry MAX_ATTEMPTS FIRST_DELAY COMMAND [ARGS...]
retry() {
local max=$1 delay=$2 attempt=1
shift 2
until "$@"; do
if (( attempt >= max )); then
echo " attempt $attempt failed, giving up"
return 1
fi
echo " attempt $attempt failed, waiting ${delay}s"
sleep "$delay"
attempt=$((attempt + 1))
delay=$((delay * 2))
done
echo " attempt $attempt worked"
}
echo "== flaky: up to 5 attempts =="
retry 5 1 flaky
echo "== always_fails: up to 2 attempts =="
retry 2 1 always_fails || echo "retry returned $?"RecapThe whole problem in a few lines, for the night before
- Spot it: "flaky command", "try again a few times"
- Idea:
until "$@"; do sleep $delay; delay=$((delay * 2)); donewith an attempt cap - Cost: the total wait grows quickly: 1 + 2 + 4 + 8 seconds
- Trap: retrying forever, or retrying errors that will never pass (404, bad password)
Interview follow-ups
Give up after a total time limit instead of an attempt count.
Record the start with
start=$SECONDS(bash's built-in seconds counter), and in the loop check(( SECONDS - start + delay > limit ))before sleeping, returning 1 if the next wait would pass the limit. You can keep the attempt limit too and stop at whichever comes first. For a hard limit on a single attempt, wrap the command itself intimeout 10, so one hung attempt cannot use all the time.
Frequently asked questions
If a thousand servers all fail at the same moment and all retry after exactly 1, 2 and 4 seconds, they hit the recovering service together each time. Jitter adds a small random amount to each wait, spreading the retries out. In bash, sleep "$(( delay + RANDOM % delay ))" waits between delay and twice that. Big cloud providers recommend backoff with jitter for exactly this reason.
Often not. curl --retry 5 --retry-delay 1 retries network errors with backoff built in, and --retry-all-errors widens it. wget --tries, package managers and many CLIs have similar options. Kubernetes, systemd (Restart=on-failure with RestartSec=) and CI systems retry at their own level. Write your own retry when a step has no built-in option, like a custom health check or a git push in CI.
It depends on how long the problem usually lasts and how long the caller can wait. For a network blip, 3 to 5 attempts starting at 1 second, doubling, covers about half a minute. Also set a maximum single wait, like 30 seconds, so the doubling does not grow to hours. For a deploy step, think about the total time: 5 attempts with doubling from 1 second can take over 30 seconds before it gives up.