Bash and Linux

Validate IPv4 Addresses

mediumgrep and regex

Problem statement

Print only the valid IPv4 addresses from a list, by building a regular expression with grep -E step by step and then checking each number with awk. Then use the same pattern to pull every IP address out of a log line. You need this when cleaning an inventory, checking a firewall allow-list, or finding which clients hit a server.

candidates.txt

TEXT
10.0.0.1
192.168.1.254
256.1.1.1
10.0.0
1.2.3.4.5
8.8.8.8
abc.def.ghi.jkl
172.16.300.1
0.0.0.0

access.log

TEXT
client 10.1.2.3 connected via gateway 10.1.2.1 port 443

An IPv4 address is four numbers from 0 to 255, joined by dots.

  1. Print the lines that have the right shape: four groups of 1 to 3 digits joined by dots, and nothing else on the line.
  2. From those, print only the ones where every number is 255 or less.
  3. Print every IP address found inside the log line, one per line.

Expected output:

TEXT
== step 1: right shape (four groups of 1 to 3 digits) ==
10.0.0.1
192.168.1.254
256.1.1.1
8.8.8.8
172.16.300.1
0.0.0.0
== step 2: right shape AND every part is 0 to 255 ==
10.0.0.1
192.168.1.254
8.8.8.8
0.0.0.0
== pull every IP out of a log line ==
10.1.2.3
10.1.2.1

Hints

Hint 1: [0-9]{1,3} means "1 to 3 digits", and \. means a real dot (a plain . means any character). ^ and $ pin the pattern to the start and end of the line.

Approach

Optimal: grep -E then awk

Covers: regular expressions, grep -E, [0-9], {1,3}, \., groups ( ), ^ and $, grep -o, awk -F. with number tests.

A regular expression describes the shape of text. Instead of searching for one exact string, you describe what the text looks like: "some digits, a dot, some digits...". grep -E turns on extended regular expressions, which let you use { }, ( ), + and | without backslashes.

Build the pattern one piece at a time.

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB A["[0-9]{1,3}
1 to 3 digits"]:::blue --> B["\.
a real dot"]:::yellow B --> C["( ... ){3}
that pair, 3 times"]:::purple C --> D["[0-9]{1,3}
one last number"]:::blue D --> E["^ and $
the whole line only"]:::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px
Piece Means Matches
[0-9] one digit 7
[0-9]{1,3} 1 to 3 digits 7, 10, 254
\. a real dot .
([0-9]{1,3}\.){3} a number and a dot, three times 10.0.0.
([0-9]{1,3}\.){3}[0-9]{1,3} then one more number 10.0.0.1
^ ... $ the whole line, nothing before or after not 1.2.3.4.5

Why the dot needs a backslash. In a regex, . means "any one character". Without the backslash, 10.0.0.1 would also match 10a0b0c1. \. means "a real dot only".

Why ^ and $ matter. Without them, grep looks for the pattern anywhere in the line. 1.2.3.4.5 contains 1.2.3.4, and 10.0.0.1abc contains 10.0.0.1, so both would pass. ^ means "start of the line" and $ means "end of the line", so the whole line must be the address.

Shape is not enough. The pattern accepts 256.1.1.1 and 172.16.300.1, because each part is 1 to 3 digits. Checking "0 to 255" inside a regex is possible, but the pattern becomes long and hard to read. It is clearer to let awk do the maths:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB IN[("9 candidates")]:::gray --> G{{"grep -E: right shape?"}}:::purple G --> X1["dropped
10.0.0, 1.2.3.4.5, abc..."]:::red G --> P["6 lines"]:::yellow P --> A{{"awk: each part 255 or less?"}}:::purple A --> X2["dropped
256.1.1.1, 172.16.300.1"]:::red A --> OK["4 valid addresses"]:::green classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px

awk -F. splits each line at the dots, so $1 to $4 are the four numbers. The test $1 <= 255 && $2 <= 255 && $3 <= 255 && $4 <= 255 keeps the line only if all four are small enough. When an awk program is just a test with no { }, it prints the lines where the test is true.

Pulling addresses out of text with -o. For a log line, you do not want the whole line, only the addresses inside it. -o (only matching) prints just the part that matched, each match on its own line. Here you drop ^ and $, because the address sits in the middle of other text.

Walking through the code. The # Setup: lines only create the two sample files, so skip past them.

  1. grep -E '^([0-9]{1,3}\.){3}[0-9]{1,3}$' keeps six lines with the right shape. 10.0.0 has only three parts, 1.2.3.4.5 has five, and abc.def.ghi.jkl has no digits, so they are dropped.
  2. The same grep is piped into awk -F., which drops 256.1.1.1 and 172.16.300.1.
  3. grep -oE without ^ and $ prints both addresses from the log line.

Edge cases. The -o step on its own would also pull 999.1.1.1 out of a log, so pipe it into the same awk check when it matters. Leading zeros like 010.0.0.1 pass both checks; some tools read them as octal, so many teams reject them. An empty file prints nothing.

# Setup: a list of candidate addresses and one log line
cd "$(mktemp -d)"
cat > candidates.txt << 'LIST'
10.0.0.1
192.168.1.254
256.1.1.1
10.0.0
1.2.3.4.5
8.8.8.8
abc.def.ghi.jkl
172.16.300.1
0.0.0.0
LIST
echo 'client 10.1.2.3 connected via gateway 10.1.2.1 port 443' > access.log

echo "== step 1: right shape (four groups of 1 to 3 digits) =="
grep -E '^([0-9]{1,3}\.){3}[0-9]{1,3}$' candidates.txt

echo "== step 2: right shape AND every part is 0 to 255 =="
grep -E '^([0-9]{1,3}\.){3}[0-9]{1,3}$' candidates.txt |
  awk -F. '$1 <= 255 && $2 <= 255 && $3 <= 255 && $4 <= 255'

echo "== pull every IP out of a log line =="
grep -oE '([0-9]{1,3}\.){3}[0-9]{1,3}' access.log

Interview follow-ups

  • Count how many times each client IP appears in a whole access log.

    Pull the addresses out, then count them: grep -oE '([0-9]{1,3}\.){3}[0-9]{1,3}' access.log | sort | uniq -c | sort -rn. sort puts equal addresses next to each other, uniq -c counts each group, and sort -rn puts the biggest count first. If the client IP is always the first field, awk '{print $1}' is simpler and faster than -o. This is the same count-and-rank pipeline as the top IPs page.

Frequently asked questions

\d (a digit) comes from Perl and other languages, and plain grep -E does not support it: it treats \d as the letter d. Use [0-9] or [[:digit:]] instead, which work everywhere. GNU grep has a -P flag for Perl-style patterns where \d does work, but -P is missing on macOS and some minimal systems. Sticking to [0-9] keeps scripts portable.

Plain grep uses basic regular expressions, where {, (, | and + must be written with a backslash to be special: \{1,3\}. grep -E uses extended regular expressions, where they are special as written, which is much easier to read. egrep is an old name for grep -E; newer GNU grep prints a warning when you use it. Use grep -E in anything new.

Yes. One part becomes (25[0-5]|2[0-4][0-9]|1[0-9]{2}|[1-9]?[0-9]): 250 to 255, or 200 to 249, or 100 to 199, or 0 to 99. Repeat that four times with dots in between and you have a full check in one pattern. It works, but it is hard to read, and a typo is hard to spot. In a script, the two-step version on this page is easier to trust and to change.