Loop Over Files Safely
Problem statement
Loop over log files in a folder and print each file's line count, skipping empty files, in a way that does not break on a name with a space or on a folder with no matches. Then do the same for subfolders with find. "For each file, do X" is in almost every script, and the unsafe versions break on the first odd file name.
The script creates these files:
files
app.log 2 linesempty.log emptymy report.log 3 lines (the name has a space)notes.txt not a logsub/deep.log 1 line, in a subfolder- Show the broken version:
for f in $(ls *.log)splitsmy report.loginto two items. - Loop over
*.logsafely, skip empty files withcontinue, and printname: N lines. - Show what an unmatched pattern like
*.csvgives, with and withoutnullglob. - Do the same over all subfolders with
find -print0andwhile read -d ''.
Expected output:
== broken: for f in $(ls *.log) ==item: app.logitem: empty.logitem: myitem: report.log4 items for 3 files== safe: for f in *.log ==app.log: 2 linesmy report.log: 3 lines== a pattern that matches nothing ==without nullglob: f is '*.csv'with nullglob: the loop ran 0 times== every .log in every subfolder ==./app.log: 2 lines./my report.log: 3 lines./sub/deep.log: 1 linesHints
for f in *.log; do ...; done, and always write "$f" inside. [[ -s $f ]] || continue skips empty files.Approach
Optimal: Globs, nullglob and find -print0
Covers: for f in *.log, quoting "$f", why for f in $(ls) breaks, shopt -s nullglob, continue and break, wc -l < file, find -print0 with while IFS= read -r -d '', process substitution < <(...).
Strict mode, used on every page in this section. The second line, set -euo pipefail, makes bash stop on mistakes instead of carrying on:
| Option | Means |
|---|---|
-e |
exit as soon as a command fails (with some exceptions, see Handle Command Failures) |
-u |
treat an unset variable as an error, instead of silently using empty text |
-o pipefail |
a pipeline fails if any command in it fails, not just the last one |
Put it right after the shebang line #!/usr/bin/env bash in every script you write.
Let the shell expand the pattern. In for f in *.log, bash itself turns *.log into the list of matching names, and each name becomes one item, spaces and all. Nothing is split. The only rule is to quote "$f" every time you use it, so the name stays one argument to wc, cp or rm.
Why for f in $(ls) breaks. $(ls *.log) gives one block of text with names separated by spaces and newlines. Unquoted, bash splits it at every space, so my report.log becomes my and report.log:
Never parse ls in scripts. Use the pattern directly, or find.
When nothing matches. By default, a pattern that matches no files is left as it is, so for f in *.csv runs once with f set to the literal text *.csv, and the next command fails with "No such file". shopt -s nullglob makes an unmatched pattern expand to nothing, so the loop simply runs zero times.
Skipping and stopping. continue jumps to the next item, break leaves the loop. [[ -s $f ]] || continue reads as "if the file is not empty, carry on; otherwise skip it".
Subfolders: find with null characters. A pattern only looks in one folder (or use shopt -s globstar and **/*.log). find walks the whole tree. To read its results safely in a loop:
| Piece | Does |
|---|---|
find ... -print0 |
ends each name with a null character |
IFS= |
keep spaces at the start and end of names |
read -r |
keep backslashes as they are |
read -d '' |
read up to the next null character, not newline |
done < <(find ...) |
feed the loop from find, without a subshell |
The < <(...) form (process substitution) matters: with find ... | while read, the loop runs in a subshell and any variables set inside are lost when it ends.
Walking through the code. The # Setup: lines only create the sample files, so skip past them.
- The broken loop counts 4 items for 3 files.
- The safe loop prints the two non-empty logs;
wc -l < "$f"prints only the number, without the name. - In an empty folder,
*.csvstays literal withoutnullglob, and gives nothing with it. - The
findloop includessub/deep.log;sort -zsorts the null-separated names so the order is fixed.
Edge cases. Names starting with - can be mistaken for options; use ./*.log or --, as in rm -- "$f". Hidden files (starting with .) are not matched by * unless you set shopt -s dotglob. A loop over thousands of files that starts a program for each one is slow; find -exec ... {} + batches them.
#!/usr/bin/env bash
set -euo pipefail
# Setup: create the sample files in a fresh temporary folder
cd "$(mktemp -d)"
printf 'a\nb\n' > app.log
: > empty.log
printf '1\n2\n3\n' > "my report.log"
echo "not a log" > notes.txt
mkdir sub && echo "x" > sub/deep.log
echo "== broken: for f in \$(ls *.log) =="
n=0
for f in $(ls *.log); do n=$((n + 1)); echo "item: $f"; done
echo "$n items for 3 files"
echo "== safe: for f in *.log =="
for f in *.log; do
[[ -s $f ]] || continue # skip empty files
echo "$f: $(wc -l < "$f") lines"
done
echo "== a pattern that matches nothing =="
mkdir none && cd none
for f in *.csv; do echo "without nullglob: f is '$f'"; done
shopt -s nullglob
count=0
for f in *.csv; do count=$((count + 1)); done
echo "with nullglob: the loop ran $count times"
shopt -u nullglob
cd ..
echo "== every .log in every subfolder =="
while IFS= read -r -d '' f; do
[[ -s $f ]] || continue
echo "$f: $(wc -l < "$f") lines"
done < <(find . -name '*.log' -print0 | sort -z)RecapThe whole problem in a few lines, for the night before
- Spot it: "for each file"
- Idea:
for f in *.log; do ... "$f"; donewithnullglob;find -print0+read -d ''for subfolders - Cost: one loop step per file
- Trap:
for f in $(ls), or forgetting quotes around"$f"
Interview follow-ups
Count the total number of lines across all logs in all subfolders.
Keep a running total in the
whileloop:total=0; while IFS= read -r -d '' f; do total=$(( total + $(wc -l < "$f") )); done < <(find . -name '*.log' -print0), thenecho "$total". Because the loop is fed with< <(...)and not a pipe,totalstill has its value afterdone. Withfind ... | while, the total would be 0 after the loop, which is a classic bug.
Frequently asked questions
The loop puts the whole name into f correctly, but every later use of an unquoted $f is split again. wc -l < $f with my report.log becomes wc -l < my report.log, which bash rejects as an "ambiguous redirect"; rm $f would try to delete my and report.log, which could be real, different files. Quoting each use keeps the name in one piece. shellcheck warns about every unquoted use.
Patterns expand in name order. For modification time, the simplest safe way is find with GNU -printf: find . -maxdepth 1 -name '*.log' -printf '%T@ %p\0' | sort -zn | cut -zd' ' -f2-, then read with while IFS= read -r -d '' f. For quick interactive work, ls -t is fine, but keep ls out of scripts. If file names are under your control and have no spaces, simpler pipelines are acceptable.
Use a pattern for files in one folder: it is fast, simple and built in. Use find when you need subfolders, or tests a pattern cannot do, like age (-mtime), size (-size) or type (-type f). bash's shopt -s globstar lets **/*.log match in subfolders too, which is handy, but it does not filter by age or size and is slow on huge trees.