Bash and Linux

Count Connections by State

mediumNetworking and connectivity

Problem statement

Read saved ss -tan output, count TCP connections in each state, find which client has the most open connections, and spot the state that points to a bug in an app. When a server runs out of connections or a service hangs, these counts tell you whether the problem is traffic, a slow peer, or an app that never closes its sockets.

ss.txt (from ss -tan)

TEXT
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 511 0.0.0.0:80 0.0.0.0:*
ESTAB 0 0 10.0.1.5:80 203.0.113.7:51544
ESTAB 0 0 10.0.1.5:80 203.0.113.7:51545
ESTAB 0 0 10.0.1.5:80 198.51.100.2:40001
ESTAB 0 0 10.0.1.5:80 203.0.113.7:51546
TIME-WAIT 0 0 10.0.1.5:80 198.51.100.9:33010
TIME-WAIT 0 0 10.0.1.5:80 198.51.100.9:33011
TIME-WAIT 0 0 10.0.1.5:80 203.0.113.8:44120
CLOSE-WAIT 1 0 10.0.1.5:8080 10.0.1.20:5432
CLOSE-WAIT 1 0 10.0.1.5:8081 10.0.1.20:5432
CLOSE-WAIT 1 0 10.0.1.5:8082 10.0.1.20:5432
SYN-SENT 0 1 10.0.1.5:8083 10.0.1.21:6379
  1. Count connections per state, most first.
  2. Print the remote IP with the most established (ESTAB) connections.
  3. Print the CLOSE-WAIT connections and what they were talking to.

Expected output:

◈ DIAGRAM
== connections per state ==
4 ESTAB
3 CLOSE-WAIT
3 TIME-WAIT
1 LISTEN
1 SYN-SENT
== remote IP with the most ESTAB connections ==
3 203.0.113.7
== CLOSE-WAIT: the app has not closed these ==
10.0.1.5:8080 -> 10.0.1.20:5432
10.0.1.5:8081 -> 10.0.1.20:5432
10.0.1.5:8082 -> 10.0.1.20:5432

Hints

Hint 1: The state is field 1 and the peer address is field 5. Skip the header and count with uniq -c or an awk array.

Approach

Optimal: ss -tan counts with awk

Covers: TCP connection states, ss -tan, ESTAB, TIME-WAIT, CLOSE-WAIT, SYN-SENT, counting with awk arrays, stripping the port, ss state filters.

A TCP connection moves through states. It opens with a handshake, carries data, and closes in a few steps. ss -tan shows every TCP socket (-a all, -n numbers) and its current state:

%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart TB S["SYN-SENT
connecting"]:::yellow --> E["ESTAB
open, working"]:::green E --> T["TIME-WAIT
we closed first"]:::blue E --> C["CLOSE-WAIT
they closed, app has not"]:::red classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px
State Means Many of them usually means
LISTEN waiting for new connections normal: one per listening port
ESTAB open and working normal traffic, or a client holding too many
TIME-WAIT closed by this side, waiting a short time before reuse normal on busy servers
CLOSE-WAIT the other side closed, this app has not a bug: the app forgot to close sockets
SYN-SENT trying to connect, no answer yet the peer is down or a firewall drops packets
SYN-RECV half-open incoming connections a flood, or a slow network

CLOSE-WAIT is the one to worry about. When the other side closes, the kernel moves the socket to CLOSE-WAIT and waits for the app to call close(). A healthy app does it quickly. A count that keeps growing means the app leaks connections, and it will eventually run out of file descriptors or connection pool slots. Restarting it clears them, but only fixing the code stops it coming back.

TIME-WAIT is usually fine. After a connection closes, the side that closed first keeps a TIME-WAIT entry for about a minute, so stray late packets are not mistaken for a new connection. Thousands of them on a busy web server are normal.

Walking through the code. The # Setup: lines only save the sample output, so skip past them.

  1. The first awk counts each state in an array, and sort -rn ranks them.
  2. The second keeps ESTAB lines, strips the port from field 5 with sub(/:[^:]*$/, "", ip), counts per IP and prints the top one.
  3. The third prints the local and peer address of each CLOSE-WAIT socket. All three are this server's app ports talking to 10.0.1.20:5432, a PostgreSQL server: the app is not closing its database connections.
%%{init: {"flowchart": {"padding": 18, "nodeSpacing": 30, "rankSpacing": 40, "htmlLabels": true}, "themeVariables": {"fontSize": "18px"}}}%% flowchart LR subgraph APP["this server's app"] direction TB P1[":8080"]:::red ~~~ P2[":8081"]:::red ~~~ P3[":8082"]:::red end APP --> DB[("10.0.1.20:5432
PostgreSQL")]:::gray classDef blue fill:#dbeafe,stroke:#2563eb,color:#1e3a8a,stroke-width:2px classDef yellow fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:2px classDef green fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:2px classDef red fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:2px classDef purple fill:#ede9fe,stroke:#7c3aed,color:#4c1d95,stroke-width:2px classDef gray fill:#f3f4f6,stroke:#6b7280,color:#111827,stroke-width:2px linkStyle default stroke:#94a3b8,stroke-width:2px style APP fill:transparent,stroke:#dc2626,stroke-width:2px

Edge cases. IPv6 peers like [2001:db8::7]:443 also lose their port with the same sub, because it removes only after the last colon; the brackets stay. On a real server, ss -tan state close-wait filters for one state directly, and ss -s prints a summary of counts.

# Setup: save sample ss -tan output in a fresh temporary folder (on a server: ss -tan)
cd "$(mktemp -d)"
cat > ss.txt << 'OUT'
State      Recv-Q Send-Q  Local Address:Port    Peer Address:Port
LISTEN     0      511           0.0.0.0:80           0.0.0.0:*
ESTAB      0      0            10.0.1.5:80        203.0.113.7:51544
ESTAB      0      0            10.0.1.5:80        203.0.113.7:51545
ESTAB      0      0            10.0.1.5:80       198.51.100.2:40001
ESTAB      0      0            10.0.1.5:80        203.0.113.7:51546
TIME-WAIT  0      0            10.0.1.5:80       198.51.100.9:33010
TIME-WAIT  0      0            10.0.1.5:80       198.51.100.9:33011
TIME-WAIT  0      0            10.0.1.5:80        203.0.113.8:44120
CLOSE-WAIT 1      0            10.0.1.5:8080       10.0.1.20:5432
CLOSE-WAIT 1      0            10.0.1.5:8081       10.0.1.20:5432
CLOSE-WAIT 1      0            10.0.1.5:8082       10.0.1.20:5432
SYN-SENT   0      1            10.0.1.5:8083       10.0.1.21:6379
OUT

echo "== connections per state =="
awk 'NR > 1 { c[$1]++ } END { for (s in c) print c[s], s }' ss.txt | sort -k1,1nr -k2,2

echo "== remote IP with the most ESTAB connections =="
awk '$1 == "ESTAB" { ip = $5; sub(/:[^:]*$/, "", ip); c[ip]++ }
     END { for (i in c) print c[i], i }' ss.txt | sort -rn | head -n 1

echo "== CLOSE-WAIT: the app has not closed these =="
awk '$1 == "CLOSE-WAIT" { print $4, "->", $5 }' ss.txt

Interview follow-ups

  • Alert when CLOSE-WAIT connections pass 100.

    Count them live: n=$(ss -tan state close-wait | tail -n +2 | wc -l), where tail -n +2 skips the header. Then [ "$n" -gt 100 ] && echo "CLOSE-WAIT: $n" && exit 1. Run it every minute from cron or a monitoring agent and graph the number, because a slow climb is the real signal. Pair the alert with the program name from ss -tanp, so whoever is paged knows which app to look at.

Frequently asked questions

Add -p and sudo: sudo ss -tanp state close-wait prints the program and PID for each socket. Count per program with sudo ss -tanp state close-wait | grep -o 'users:(("[^"]*' | sort | uniq -c. Then look at that app's connection handling, often a database or HTTP client pool that is not returning connections. A growing count over time, not one snapshot, is what proves the leak.

Usually no. TIME-WAIT entries are cheap and protect against late packets. They only cause trouble when a client makes many short connections to the same server and port and runs out of local ports. The better fixes are reusing connections (HTTP keep-alive, connection pools) and widening the port range (net.ipv4.ip_local_port_range). Old advice to enable tcp_tw_recycle is dangerous, and that option has been removed from modern kernels.

Recv-Q is data that has arrived but the app has not read yet; if it grows on established connections, the app is too slow or stuck. Send-Q is data sent but not yet acknowledged by the other side; if it grows, the network or the peer is slow. On LISTEN sockets, Recv-Q is the number of connections waiting to be accepted and Send-Q is the backlog limit, so a Recv-Q near the limit means new connections are being dropped.