Bash and Linux

Loop Over Files Safely

easyBash scripts and automation Must-do

Problem statement

Loop over log files in a folder and print each file's line count, skipping empty files, in a way that does not break on a name with a space or on a folder with no matches. Then do the same for subfolders with find. "For each file, do X" is in almost every script, and the unsafe versions break on the first odd file name.

The script creates these files:

files

TEXT
app.log 2 lines
empty.log empty
my report.log 3 lines (the name has a space)
notes.txt not a log
sub/deep.log 1 line, in a subfolder
  1. Show the broken version: for f in $(ls *.log) splits my report.log into two items.
  2. Loop over *.log safely, skip empty files with continue, and print name: N lines.
  3. Show what an unmatched pattern like *.csv gives, with and without nullglob.
  4. Do the same over all subfolders with find -print0 and while read -d ''.

Expected output:

TEXT
== broken: for f in $(ls *.log) ==
item: app.log
item: empty.log
item: my
item: report.log
4 items for 3 files
== safe: for f in *.log ==
app.log: 2 lines
my report.log: 3 lines
== a pattern that matches nothing ==
without nullglob: f is '*.csv'
with nullglob: the loop ran 0 times
== every .log in every subfolder ==
./app.log: 2 lines
./my report.log: 3 lines
./sub/deep.log: 1 lines

Hints

Hint 1: Let the shell expand the pattern itself: for f in *.log; do ...; done, and always write "$f" inside. [[ -s $f ]] || continue skips empty files.

Approach

Optimal: Globs, nullglob and find -print0

Covers: for f in *.log, quoting "$f", why for f in $(ls) breaks, shopt -s nullglob, continue and break, wc -l < file, find -print0 with while IFS= read -r -d '', process substitution < <(...).

Strict mode, used on every page in this section. The second line, set -euo pipefail, makes bash stop on mistakes instead of carrying on:

Option Means
-e exit as soon as a command fails (with some exceptions, see Handle Command Failures)
-u treat an unset variable as an error, instead of silently using empty text
-o pipefail a pipeline fails if any command in it fails, not just the last one

Put it right after the shebang line #!/usr/bin/env bash in every script you write.

Let the shell expand the pattern. In for f in *.log, bash itself turns *.log into the list of matching names, and each name becomes one item, spaces and all. Nothing is split. The only rule is to quote "$f" every time you use it, so the name stays one argument to wc, cp or rm.

Why for f in $(ls) breaks. $(ls *.log) gives one block of text with names separated by spaces and newlines. Unquoted, bash splits it at every space, so my report.log becomes my and report.log:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart LR subgraph BAD["for f in $(ls *.log)"] direction TB B1["app.log"]:::gray ~~~ B2["empty.log"]:::gray ~~~ B3["my"]:::red ~~~ B4["report.log"]:::red end subgraph GOOD["for f in *.log"] direction TB G1["app.log"]:::gray ~~~ G2["empty.log"]:::gray ~~~ G3["my report.log"]:::green end BAD ~~~ GOOD classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px style BAD fill:transparent,stroke:#dc2626,stroke-width:2px style GOOD fill:transparent,stroke:#059669,stroke-width:2px

Never parse ls in scripts. Use the pattern directly, or find.

When nothing matches. By default, a pattern that matches no files is left as it is, so for f in *.csv runs once with f set to the literal text *.csv, and the next command fails with "No such file". shopt -s nullglob makes an unmatched pattern expand to nothing, so the loop simply runs zero times.

Skipping and stopping. continue jumps to the next item, break leaves the loop. [[ -s $f ]] || continue reads as "if the file is not empty, carry on; otherwise skip it".

Subfolders: find with null characters. A pattern only looks in one folder (or use shopt -s globstar and **/*.log). find walks the whole tree. To read its results safely in a loop:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB F(["find . -name '*.log' -print0"]):::purple --> R(["while IFS= read -r -d '' f"]):::purple R --> S{{"-s "$f"?"}}:::yellow S --> P["print name and lines"]:::green S --> C["continue"]:::gray classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px
Piece Does
find ... -print0 ends each name with a null character
IFS= keep spaces at the start and end of names
read -r keep backslashes as they are
read -d '' read up to the next null character, not newline
done < <(find ...) feed the loop from find, without a subshell

The < <(...) form (process substitution) matters: with find ... | while read, the loop runs in a subshell and any variables set inside are lost when it ends.

Walking through the code. The # Setup: lines only create the sample files, so skip past them.

  1. The broken loop counts 4 items for 3 files.
  2. The safe loop prints the two non-empty logs; wc -l < "$f" prints only the number, without the name.
  3. In an empty folder, *.csv stays literal without nullglob, and gives nothing with it.
  4. The find loop includes sub/deep.log; sort -z sorts the null-separated names so the order is fixed.

Edge cases. Names starting with - can be mistaken for options; use ./*.log or --, as in rm -- "$f". Hidden files (starting with .) are not matched by * unless you set shopt -s dotglob. A loop over thousands of files that starts a program for each one is slow; find -exec ... {} + batches them.

#!/usr/bin/env bash
set -euo pipefail

# Setup: create the sample files in a fresh temporary folder
cd "$(mktemp -d)"
printf 'a\nb\n' > app.log
: > empty.log
printf '1\n2\n3\n' > "my report.log"
echo "not a log" > notes.txt
mkdir sub && echo "x" > sub/deep.log

echo "== broken: for f in \$(ls *.log) =="
n=0
for f in $(ls *.log); do n=$((n + 1)); echo "item: $f"; done
echo "$n items for 3 files"

echo "== safe: for f in *.log =="
for f in *.log; do
  [[ -s $f ]] || continue            # skip empty files
  echo "$f: $(wc -l < "$f") lines"
done

echo "== a pattern that matches nothing =="
mkdir none && cd none
for f in *.csv; do echo "without nullglob: f is '$f'"; done
shopt -s nullglob
count=0
for f in *.csv; do count=$((count + 1)); done
echo "with nullglob: the loop ran $count times"
shopt -u nullglob
cd ..

echo "== every .log in every subfolder =="
while IFS= read -r -d '' f; do
  [[ -s $f ]] || continue
  echo "$f: $(wc -l < "$f") lines"
done < <(find . -name '*.log' -print0 | sort -z)
RecapThe whole problem in a few lines, for the night before
  • Spot it: "for each file"
  • Idea: for f in *.log; do ... "$f"; done with nullglob; find -print0 + read -d '' for subfolders
  • Cost: one loop step per file
  • Trap: for f in $(ls), or forgetting quotes around "$f"

Interview follow-ups

  • Count the total number of lines across all logs in all subfolders.

    Keep a running total in the while loop: total=0; while IFS= read -r -d '' f; do total=$(( total + $(wc -l < "$f") )); done < <(find . -name '*.log' -print0), then echo "$total". Because the loop is fed with < <(...) and not a pipe, total still has its value after done. With find ... | while, the total would be 0 after the loop, which is a classic bug.

Frequently asked questions

The loop puts the whole name into f correctly, but every later use of an unquoted $f is split again. wc -l < $f with my report.log becomes wc -l < my report.log, which bash rejects as an "ambiguous redirect"; rm $f would try to delete my and report.log, which could be real, different files. Quoting each use keeps the name in one piece. shellcheck warns about every unquoted use.

Patterns expand in name order. For modification time, the simplest safe way is find with GNU -printf: find . -maxdepth 1 -name '*.log' -printf '%T@ %p\0' | sort -zn | cut -zd' ' -f2-, then read with while IFS= read -r -d '' f. For quick interactive work, ls -t is fine, but keep ls out of scripts. If file names are under your control and have no spaces, simpler pipelines are acceptable.

Use a pattern for files in one folder: it is fast, simple and built in. Use find when you need subfolders, or tests a pattern cannot do, like age (-mtime), size (-size) or type (-type f). bash's shopt -s globstar lets **/*.log match in subfolders too, which is handy, but it does not filter by age or size and is slow on huge trees.