Search Rotated Logs
Problem statement
Search a log together with its rotated copies, some of them compressed, and read the matches in time order. When an incident started yesterday, the evidence is in app.log.1 or app.log.2.gz, not in app.log, and plain grep cannot see inside .gz files.
The script creates a log and three rotated copies the way logrotate leaves them (newest is app.log, oldest is app.log.3.gz):
files
app.log today 3 lines, 1 timeoutapp.log.1 yesterday 3 lines, 1 timeoutapp.log.2.gz 2 days ago 3 lines, 2 timeouts (compressed)app.log.3.gz 3 days ago 2 lines, 0 timeouts (compressed)- Show that plain
grepfinds nothing inside a.gzfile. - Count the
timeoutlines in every file withzgrep. - Print every
timeoutline from oldest to newest, as one stream.
Expected output:
== plain grep cannot see inside a .gz file ==0== zgrep: matches per file ==app.log:1app.log.1:1app.log.2.gz:2app.log.3.gz:0== every timeout, oldest first ==2026-09-29T09:10:00Z ERROR db timeout2026-09-29T09:20:00Z ERROR db timeout2026-09-30T14:00:00Z ERROR api timeout2026-10-01T09:30:00Z ERROR db timeoutHints
zgrep works like grep but opens .gz files first; it reads plain files too. zgrep -c timeout app.log* counts per file.Approach
Optimal: zgrep and zcat -f, oldest first
Covers: log rotation, .1 and .gz naming, gzip, zcat, zcat -f, zgrep, zless, globs like app.log*, putting files in time order, ls -tr.
What rotation does. A log cannot grow forever, so logrotate (or the app itself) regularly renames it and starts a fresh file. Older copies get higher numbers and are usually compressed to save space:
today"]:::green ~~~ B["app.log.1
yesterday"]:::blue ~~~ C[("app.log.2.gz
2 days, compressed")]:::yellow ~~~ D[("app.log.3.gz
3 days, compressed")]:::gray end classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px style ROT fill:transparent,stroke:#6b7280,stroke-width:2px
| File | Age | Compressed |
|---|---|---|
app.log |
now, still being written | no |
app.log.1 |
the last rotation | usually not (delaycompress) |
app.log.2.gz |
the one before | yes |
app.log.3.gz |
older still | yes |
Higher number = older. A glob like app.log* matches all of them.
grep cannot read compressed files. A .gz file is compressed binary data, so the word timeout does not appear in it as text. grep -c timeout app.log.2.gz prints 0 even though there are two matches inside. This is the classic way to miss the evidence.
The z tools uncompress on the fly.
| Command | Does |
|---|---|
zcat file.gz |
print it uncompressed |
zcat -f file |
the same, and plain files pass through as they are |
zgrep pattern files |
grep inside plain and .gz files |
zless file.gz |
page through it |
gzip -dc file.gz |
the same as zcat |
Nothing is uncompressed to disk, so this is safe on a full server.
Reading in time order. zgrep searches files in the order you list them, and the glob lists app.log (newest) first. For a timeline, list the files oldest first and join them into one stream: zcat -f app.log.3.gz app.log.2.gz app.log.1 app.log | grep timeout. Now the matches read from oldest to newest, with no file names in the way.
.3.gz .2.gz .1 app.log"]:::blue --> Z(["zcat -f
one stream of text"]):::purple Z --> G(["grep timeout"]):::purple G --> T["matches in time order"]:::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px
Walking through the code. The # Setup: lines only write four small logs and compress two of them with gzip -n, so skip past them.
- Plain
grep -con the.gzfile prints 0.|| truekeeps the script going, because grep exits with 1 when it finds nothing. zgrep -c timeout app.log*printsfile:countfor every file, compressed or not.zcat -fjoins all four, oldest first, andgrepprints the four timeout lines in time order.
Edge cases. Some setups use other compressors, like .bz2 (bzgrep, bzcat), .xz (xzgrep, xzcat) or .zst (zstdgrep, zstdcat). Date-named files like app.log-20260930.gz sort by name in time order, which makes the oldest-first list easy: ls app.log-*. With numbered names, a plain sort puts .10 before .2; use ls -tr (oldest modification first) or sort -V.
# Setup: a log and three rotated copies, two of them gzip-compressed, in a fresh temporary folder
cd "$(mktemp -d)"
printf '%s\n' '2026-09-28T08:00:00Z INFO start' '2026-09-28T09:00:00Z INFO ok' > app.log.3
printf '%s\n' '2026-09-29T08:00:00Z INFO start' '2026-09-29T09:10:00Z ERROR db timeout' \
'2026-09-29T09:20:00Z ERROR db timeout' > app.log.2
printf '%s\n' '2026-09-30T08:00:00Z INFO start' '2026-09-30T14:00:00Z ERROR api timeout' \
'2026-09-30T15:00:00Z INFO ok' > app.log.1
printf '%s\n' '2026-10-01T08:00:00Z INFO start' '2026-10-01T09:30:00Z ERROR db timeout' \
'2026-10-01T10:00:00Z INFO ok' > app.log
gzip -n app.log.3 app.log.2
echo "== plain grep cannot see inside a .gz file =="
grep -c timeout app.log.2.gz || true
echo "== zgrep: matches per file =="
zgrep -c timeout app.log*
echo "== every timeout, oldest first =="
zcat -f app.log.3.gz app.log.2.gz app.log.1 app.log | grep timeoutRecapThe whole problem in a few lines, for the night before
- Spot it: "search yesterday's logs too", files ending in
.1and.gz - Idea:
zgrep pattern app.log*per file;zcat -f oldest ... newest | grepfor a timeline - Cost: uncompresses in memory as it reads, nothing written to disk
- Trap: plain
grepon.gzfiles finds nothing, andapp.log*lists newest first
Interview follow-ups
Print each match with the name of the file it came from, in time order.
Loop over the files oldest first and add the name in front of each match:
for f in app.log.3.gz app.log.2.gz app.log.1 app.log; do zgrep -H timeout "$f"; done.-Hmakes grep print the file name even for a single file. Because the loop controls the order, the output is a timeline with sources. In GNU grep you can add--labelwhen reading from a pipe, but the loop is clearer.
Frequently asked questions
Classic text logs live in /var/log (syslog or messages, auth.log or secure, plus folders for nginx, apt and others) and are rotated by logrotate, with the config in /etc/logrotate.conf and /etc/logrotate.d/. On systemd systems, much of the same data is also in the journal, which rotates itself by size and age; search it with journalctl --since "2 days ago" | grep timeout. Containers usually log to stdout, read with docker logs or kubectl logs --previous.
sudo zgrep -l 'pattern' /var/log/*.log /var/log/*.gz lists the files that contain it. For subfolders, combine with find: sudo find /var/log -type f \( -name '*.log*' \) -exec zgrep -l 'pattern' {} +. Some files in /var/log, like wtmp and lastlog, are binary records, not text; read them with last and lastlog instead of grep.
Narrow the files first: only search the days that matter instead of every rotation. Use LC_ALL=C zgrep -F 'fixed text', which skips locale handling and pattern logic. Tools like ripgrep (rg -z) search compressed files with several CPU cores and are much faster. If you search the same logs often, that is a sign they should be in a log platform with an index.