Bash and Linux

Read a File Line by Line

easyBash scripts and automation Must-do

Problem statement

Read a .env file line by line, skip comments and blank lines, load each KEY=VALUE pair, and report a malformed line, without losing the last line or your variables. Reading config files, host lists and CSVs one line at a time is one of the most common jobs a script does.

.env (the last line has no newline at the end, as often happens)

Bash
# app settings
APP_NAME=shop
PORT=8080
GREETING=hello world
# indented comment
BAD LINE
LAST_KEY=no newline at the end
  1. Count lines with a plain while read, then with the fix for a missing final newline.
  2. Show that a counter set inside cat file | while is lost after the loop.
  3. Load the settings, print each one, report the bad line with its line number, and print a summary.

Expected output:

Bash
== counting lines ==
plain while read: 7 lines (last one lost)
with || [[ -n $line ]]: 8 lines
== cat | while loses the counter ==
after the loop: n=0
== loading settings ==
APP_NAME = shop
PORT = 8080
GREETING = hello world
line 7: not KEY=VALUE: BAD LINE
LAST_KEY = no newline at the end
loaded 4 settings, 1 bad line(s); PORT is 8080

Hints

Hint 1: The safe loop is while IFS= read -r line || [[ -n $line ]]; do ...; done < .env. The || [[ -n $line ]] part keeps a last line that has no newline.

Approach

Optimal: while IFS= read -r with < file

Covers: while IFS= read -r line, why IFS= and -r, the redirect after done, the last-line-without-newline fix, cat file | while and subshells, =~ and BASH_REMATCH, associative arrays for settings, why not for line in $(cat file).

Strict mode, used on every page in this section. The second line, set -euo pipefail, makes bash stop on mistakes instead of carrying on:

Option Means
-e exit as soon as a command fails (with some exceptions, see Handle Command Failures)
-u treat an unset variable as an error, instead of silently using empty text
-o pipefail a pipeline fails if any command in it fails, not just the last one

Put it right after the shebang line #!/usr/bin/env bash in every script you write.

The standard loop. This exact line is worth memorising:

TEXT
while IFS= read -r line || [[ -n $line ]]; do
...
done < file
%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB F[(".env")]:::blue --> R(["IFS= read -r line"]):::purple R --> B["loop body
uses "$line""]:::green B --> R R --> E["end of file: loop ends"]:::gray classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px
Piece Why
read line reads one line into line; fails at the end of the file, which ends the loop
IFS= keep spaces at the start and end of the line (they are trimmed otherwise)
-r keep backslashes as they are, instead of treating them as escapes
|| [[ -n $line ]] read fails on a last line with no newline, but still fills line; this keeps it
done < file the file feeds the whole loop, in the current shell

Why not cat file | while? Each part of a pipeline runs in a subshell, a copy of the shell. The loop then counts in the copy, and when the pipe ends the copy disappears with its variables. After the loop, the counter is back at 0. done < file keeps the loop in your shell, so variables survive.

Why not for line in $(cat file)? That splits the file at every space, not every line, so GREETING=hello world becomes two items, and it also expands * into file names. for is for lists of words; while read is for lines.

Recognising each kind of line. For every line, the loop decides:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB L["next line"]:::gray --> C{{"comment or blank?"}}:::yellow C --> S["skip"]:::gray C --> K{{"KEY=VALUE?"}}:::purple K --> OK["store in config"]:::green K --> BAD["report the line number"]:::red classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px

[[ $line =~ regex ]] tests a regular expression, and each ( ) group's match is stored in BASH_REMATCH. The key pattern [A-Za-z_][A-Za-z0-9_]* is a valid variable name; everything after the first = is the value, spaces included. The settings go into an associative array, so they can be looked up by name later.

Walking through the code. The # Setup: lines only write .env, with printf and no final newline on the last line, so skip past them.

  1. The plain loop counts 7 lines and misses LAST_KEY; the fixed loop counts 8.
  2. The cat | while version ends with the counter still 0.
  3. The parser keeps a line number, skips comments and blanks, loads 4 settings, reports line 7, and prints PORT from the array.

The counter uses n=$((n + 1)), not ((n++)), because ((n++)) returns failure when n is 0 and set -e would stop the script.

Edge cases. Values with quotes, like NAME="a b", keep their quotes here; strip them if your format uses them. Do not source a .env you do not control, because sourcing runs any command written in it. Windows line endings leave \r at the end of every value; remove it with line=${line%$'\r'}.

#!/usr/bin/env bash
set -euo pipefail

# Setup: write the sample .env (no newline after the last line) in a fresh temporary folder
cd "$(mktemp -d)"
printf '%s\n' '# app settings' 'APP_NAME=shop' '' 'PORT=8080' 'GREETING=hello world' \
              '  # indented comment' 'BAD LINE' > .env
printf 'LAST_KEY=no newline at the end' >> .env

echo "== counting lines =="
n=0; while IFS= read -r line; do n=$((n + 1)); done < .env
echo "plain while read:        $n lines (last one lost)"
n=0; while IFS= read -r line || [[ -n $line ]]; do n=$((n + 1)); done < .env
echo "with || [[ -n \$line ]]: $n lines"

echo "== cat | while loses the counter =="
n=0; cat .env | while IFS= read -r line || [[ -n $line ]]; do n=$((n + 1)); done
echo "after the loop: n=$n"

echo "== loading settings =="
declare -A config
lineno=0 bad=0
re='^([A-Za-z_][A-Za-z0-9_]*)=(.*)$'
while IFS= read -r line || [[ -n $line ]]; do
  lineno=$((lineno + 1))
  [[ $line =~ ^[[:space:]]*(#|$) ]] && continue        # comment or blank
  if [[ $line =~ $re ]]; then
    config[${BASH_REMATCH[1]}]=${BASH_REMATCH[2]}
    echo "${BASH_REMATCH[1]} = ${BASH_REMATCH[2]}"
  else
    echo "line $lineno: not KEY=VALUE: $line"
    bad=$((bad + 1))
  fi
done < .env
echo "loaded ${#config[@]} settings, $bad bad line(s); PORT is ${config[PORT]}"
RecapThe whole problem in a few lines, for the night before
  • Spot it: "for each line in a file"
  • Idea: while IFS= read -r line || [[ -n $line ]]; do ...; done < file
  • Cost: one pass; fine for config sizes, slow for millions of lines (use awk)
  • Trap: cat file | while losing variables, or for line in $(cat file) splitting at spaces

Interview follow-ups

  • Read a user;password file and print each user, without ever printing the password.

    Use while IFS=';' read -r user pass || [[ -n $user ]]; do ...; done < users.txt. Skip blank lines with [[ -n $user ]] || continue, and check both fields are present with [[ -n $pass ]]. Print only user and, if you must say something about the password, its length, ${#pass}. Never echo the password or pass it as a command-line argument, which other users can see in ps; tools like chpasswd read user:password from stdin for exactly this reason.

Frequently asked questions

read returns failure when it reaches the end of the file without finding a newline, even though it did read the text into the variable. Many editors and tools write files without a final newline, so the last line is silently lost. Adding || [[ -n $line ]] to the loop condition keeps going for that one last partial line. Writing files with printf '%s\n' or echo always adds the newline, which avoids the problem on your side.

Every part of a pipeline runs in its own subshell, and a subshell cannot change the parent's variables. The loop's changes vanish when the pipe ends. Feed the loop with a redirect instead, done < file, or from a command with done < <(command). In bash 4.2+, shopt -s lastpipe also runs the last part of a pipe in the current shell, but only when job control is off, as in scripts.

Set IFS to the separator for the read alone, and give read several names: while IFS=';' read -r user pass; do ...; done < users.txt. Each field goes into its own variable, and the last variable gets the rest of the line. For CSV with quoted commas inside fields, use a real CSV tool, because read cannot handle quotes.