Skip to main content

Networking for SRE

Learn to debug production networks like an SRE: TCP states and queues, DNS, TLS, HTTP versions, Kubernetes service paths, and load balancer behaviour.

~3.5 hours
13 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Networking Breaks Services Nobody Touched

When nothing was deployed and the service is still failing, the network is the first suspect, and the skill is proving which part of it.

Reading the TCP State Machine Under Load

A TCP connection is more than open or closed, and the two states that cause the most real outages are TIME_WAIT and CLOSE_WAIT.

Tuning TCP Queues, Buffers, and Keepalive

TIME_WAIT and CLOSE_WAIT describe the end of a connection; this topic covers its birth, where a full queue makes clients hang with no error on the...

Debugging DNS the Way Production Actually Fails

DNS rarely fails outright; it fails for some users, some of the time, usually when traffic spikes.

Diagnosing TLS Handshake Failures

HTTPS failing is a common emergency, and it takes minutes to diagnose once you know which step of the handshake broke.

Comparing HTTP/1.1, HTTP/2, and HTTP/3

Each HTTP version fixed the last one's biggest bottleneck and moved the problem somewhere new, which changes how you read a "slow request" report.

Skills You'll Master

NETWORKINGTCPDNSTLSSRE

Curriculum Index13 topics

1

Understanding Why Networking Breaks Services Nobody Touched

When nothing was deployed and the service is still failing, the network is the first suspect, and the skill is proving...

2

Reading the TCP State Machine Under Load

A TCP connection is more than open or closed, and the two states that cause the most real outages are TIME_WAIT and...

3

Tuning TCP Queues, Buffers, and Keepalive

TIME_WAIT and CLOSE_WAIT describe the end of a connection; this topic covers its birth, where a full queue makes...

4

Debugging DNS the Way Production Actually Fails

DNS rarely fails outright; it fails for some users, some of the time, usually when traffic spikes.

5

Diagnosing TLS Handshake Failures

HTTPS failing is a common emergency, and it takes minutes to diagnose once you know which step of the handshake broke.

6

Comparing HTTP/1.1, HTTP/2, and HTTP/3

Each HTTP version fixed the last one's biggest bottleneck and moved the problem somewhere new, which changes how you...

7

Finding Where Firewalls Silently Drop Packets

Most "connection timed out" incidents are a firewall rule that drops packets without telling anyone.

8

Following a Packet Through Kubernetes Services

Inside a cluster every host-to-host call is a chain of virtual interfaces and kernel rules, and most "pod cannot reach...

9

Probing gRPC Services Correctly

Standard probes cannot see inside gRPC, which is why gRPC deployments so often flap in and out of rotation or stay...

10

Choosing Load Balancer Algorithms

Every algorithm has a traffic shape where it makes latency worse, and picking the wrong one is a quiet cause of one...

11

Using the Diagnostic Toolkit

Four tools answer almost every network question: ss for sockets, tcpdump for packets, curl for end-to-end timing, and...

12

Hands-on Lab: Break and Fix the Network at acme-shop

📌 Remember: This lab is free. Part A runs on sre-vm, a local Ubuntu 24.04 virtual machine, so nothing touches your...

13

Quick Reference and Common Mistakes

Using tcp_tw_recycle to fix TIME_WAIT. Old tutorials still recommend it because it sounds like exactly the cure.

Career Impact

Roles that use the skills in this module.

  • Site Reliability Engineer

  • DevOps Engineer

  • Platform Engineer

  • Cloud Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

TIME_WAIT appears on the side that closed the connection first and is normal; it only hurts when thousands of short connections use up your ports. CLOSE_WAIT appears on the side that received the close and is waiting for its own application to call close(). A growing CLOSE_WAIT count is almost always a bug in that application.

When the accept queue is full, the kernel quietly ignores new connection attempts instead of refusing them. The client waits and times out, and the server's logs show nothing. You find it by comparing Recv-Q and Send-Q in ss for the listening socket.

Pods default to ndots:5, so a name such as example.com is tried against several cluster search domains before the real lookup. That multiplies queries and adds latency, especially when CoreDNS is CPU-throttled. Using a trailing dot, lowering ndots, or running NodeLocal DNSCache reduces it.

A TCP probe only proves the port accepts connections, and gRPC health lives inside an RPC response. A server with an exhausted database pool still accepts connections. Use the gRPC health checking protocol with Kubernetes' native gRPC probe.