Bash and Linux

Errors per Hour

mediumLogs and troubleshooting Must-do

Problem statement

Count the ERROR lines in an application log for each hour and draw a small text bar chart, so a spike stands out at a glance. "When did the errors start, and when did they stop?" is the first question in every incident review, and bucketing by time is how you answer it.

app.log

TEXT
2026-10-01T09:12:03Z INFO order created id=101
2026-10-01T09:20:11Z ERROR payment timeout id=102
2026-10-01T09:48:40Z ERROR payment timeout id=103
2026-10-01T10:01:02Z WARN slow query 2.1s
2026-10-01T10:05:17Z ERROR db connection refused
2026-10-01T10:05:19Z ERROR db connection refused
2026-10-01T10:06:00Z ERROR db connection refused
2026-10-01T10:30:44Z ERROR payment timeout id=110
2026-10-01T10:59:59Z ERROR cache miss storm
2026-10-01T11:15:00Z INFO recovered
2026-10-01T11:40:21Z ERROR payment timeout id=120
  1. Print the number of errors per hour, in time order, with a bar of # signs.
  2. Print the hour with the most errors.
  3. Print the most common error messages, ignoring the changing id= values.

Expected output:

TEXT
== errors per hour ==
2026-10-01T09:00 2 ##
2026-10-01T10:00 5 #####
2026-10-01T11:00 1 #
== busiest hour ==
2026-10-01T10:00 5
== most common error messages ==
4 payment timeout
3 db connection refused
1 cache miss storm

Hints

Hint 1: The hour is the first 13 characters of the timestamp, like 2026-10-01T10. substr($1, 1, 13) cuts it out, and an awk array counts errors per hour.

Approach

Optimal: Bucket by time with substr

Covers: time buckets, substr on ISO timestamps, awk arrays, sorting by time, text bar charts, normalising messages before counting, uniq -c.

A bucket is a slice of time. To see when errors happened, you cut each timestamp down to the size of the bucket you want, and count per bucket:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB T["2026-10-01T10:05:17Z ERROR ..."]:::gray --> C(["substr($1, 1, 13)"]):::purple C --> B["bucket
2026-10-01T10"]:::blue B --> N["c[bucket] += 1"]:::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px
Bucket Keep From 2026-10-01T10:05:17Z
day first 10 characters 2026-10-01
hour first 13 2026-10-01T10
10 minutes first 15 2026-10-01T10:0
minute first 16 2026-10-01T10:05

ISO 8601 timestamps (YYYY-MM-DDTHH:MM:SS) put the biggest unit first, so cutting the end off always gives a bigger bucket, and plain text sorting puts them in time order. That is a big reason to log in ISO format and UTC.

Counting per bucket. $2 == "ERROR" { c[substr($1, 1, 13)]++ } adds one to the box for that hour. In END, awk prints each hour and its count. The order of for (h in c) is not fixed, so the output goes through sort, which puts the hours in time order.

A bar makes the spike visible. Printing a # for each error turns the numbers into a tiny chart. The second awk builds the bar with a loop, one # per error. The 10 hour clearly stands out.

Group messages, not lines. Error messages often contain values that change, like IDs, times or user names. Counted as they are, every line is different. Removing the changing part first, here id=102, turns them into the same message:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart LR subgraph RAW["raw messages"] direction TB R1["payment timeout id=102"]:::red ~~~ R2["payment timeout id=103"]:::red ~~~ R3["payment timeout id=110"]:::red end RAW --> ONE["payment timeout
x 4"]:::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px style RAW fill:transparent,stroke:#dc2626,stroke-width:2px

Walking through the code. The # Setup: lines only create the sample log, so skip past them.

  1. The first awk counts errors per hour; sort puts the hours in order; the second awk prints each hour with its count and bar.
  2. The same counts sorted by number with sort -k2,2nr give the busiest hour.
  3. The third awk deletes $1 and $2 (the time and level), removes id= and a number, and prints the rest. sort | uniq -c | sort -rn counts the messages.

Edge cases. Hours with no errors do not appear at all; for a chart with gaps filled in, loop over the expected hours in END. Logs in other formats need a different cut: for Oct 1 10:05:17, use the third field and substr($3, 1, 2). Logs in local time can show a gap or a double hour when clocks change.

# Setup: create the sample log in a fresh temporary folder
cd "$(mktemp -d)"
cat > app.log << 'LOG'
2026-10-01T09:12:03Z INFO  order created id=101
2026-10-01T09:20:11Z ERROR payment timeout id=102
2026-10-01T09:48:40Z ERROR payment timeout id=103
2026-10-01T10:01:02Z WARN  slow query 2.1s
2026-10-01T10:05:17Z ERROR db connection refused
2026-10-01T10:05:19Z ERROR db connection refused
2026-10-01T10:06:00Z ERROR db connection refused
2026-10-01T10:30:44Z ERROR payment timeout id=110
2026-10-01T10:59:59Z ERROR cache miss storm
2026-10-01T11:15:00Z INFO  recovered
2026-10-01T11:40:21Z ERROR payment timeout id=120
LOG

echo "== errors per hour =="
awk '$2 == "ERROR" { c[substr($1, 1, 13)]++ } END { for (h in c) print h, c[h] }' app.log | sort |
  awk '{ bar = ""; for (i = 0; i < $2; i++) bar = bar "#"; printf "%s:00  %2d  %s\n", $1, $2, bar }'

echo "== busiest hour =="
awk '$2 == "ERROR" { c[substr($1, 1, 13)]++ } END { for (h in c) print h ":00", c[h] }' app.log |
  sort -k2,2nr | head -n 1

echo "== most common error messages =="
awk '$2 == "ERROR" { $1 = ""; $2 = ""; sub(/ id=[0-9]+/, ""); sub(/^ +/, ""); print }' app.log |
  sort | uniq -c | sort -k1,1nr -k2
RecapThe whole problem in a few lines, for the night before
  • Spot it: "when did errors spike", "errors per hour or minute"
  • Idea: awk '$2 == "ERROR" { c[substr($1, 1, 13)]++ }', then sort by time
  • Cost: one pass, one counter per bucket
  • Trap: counting raw messages with changing IDs, or forgetting that awk array order is random

Interview follow-ups

  • Show errors per 10 minutes for the busy hour only.

    Filter to that hour, then use a 15-character cut, which keeps the tens digit of the minutes: awk '$2 == "ERROR" && substr($1, 1, 13) == "2026-10-01T10" { c[substr($1, 1, 15)]++ } END { for (b in c) print b "0", c[b] }' app.log | sort. Adding "0" turns 10:0 back into a readable 10:00. Zooming in step by step, day to hour to minutes, is how you find the exact moment a problem began.

Frequently asked questions

Decide the range first, then loop over every bucket and print 0 where there is no count. With GNU date: for h in $(seq 0 23); do printf '2026-10-01T%02d\n' "$h"; done makes the list of hours, and join -a1 -e0 -o 0,2.2 (on sorted input) attaches the counts with 0 for missing ones. In awk, you can loop for (h = 0; h < 24; h++) in END and look up c[sprintf("2026-10-01T%02d", h)] + 0. Gaps matter, because a silent hour may mean the app was down, not healthy.

Because a single problem can produce thousands of lines that differ only in an ID, a time or an amount. Counted raw, each line is unique and the top list is useless. Replacing the changing parts with nothing, or with a placeholder like id=N, groups them into one message with a big count. Common patterns to strip are numbers ([0-9]+), UUIDs, IP addresses and quoted values.

For production systems, yes: tools like Loki, Elasticsearch, CloudWatch Logs Insights or Datadog do time buckets and graphs over all servers at once. The shell version is still worth knowing for the cases that matter most: when the log platform is down, when you are on a single server, when logs were never shipped, or when you are handed a log file in a ticket. The same idea, bucket then count, is what those platforms run underneath.