Skip to main content

Python for Production SRE Work

Learn to write Python that survives production: safe scripts, resilient API calls, tests, and CLI tools that SRE teams rely on during incidents.

~3.5 hours
13 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why SREs Need Production-Grade Python

It is the morning after acme-shop's sale-day incident. Your lead asks a simple question: how much error budget does checkout-service have left?

Writing Scripts That Fail Safely

Python basics you will use in every script Operational scripts lean on a small set of Python building blocks: strings, numbers, lists, dictionaries...

Making Script Output Useful to Humans and Machines

Replacing print with logging The logging module adds a timestamp and a severity level to every message, which print() cannot do.

Running System Commands Safely

Calling commands with subprocess The subprocess module runs shell commands from Python.

Calling APIs Safely

Making a request and checking the result The requests library is the standard way to call an HTTP API.

Making Network Calls Resilient

Setting a timeout on every call By default, requests waits forever for a response.

Skills You'll Master

PYTHONSREAUTOMATIONCLI-TOOLSTESTING

Curriculum Index13 topics

1

Understanding Why SREs Need Production-Grade Python

It is the morning after acme-shop's sale-day incident.

2

Writing Scripts That Fail Safely

Python basics you will use in every script Operational scripts lean on a small set of Python building blocks: strings...

3

Making Script Output Useful to Humans and Machines

Replacing print with logging The logging module adds a timestamp and a severity level to every message, which print()...

4

Running System Commands Safely

Calling commands with subprocess The subprocess module runs shell commands from Python.

5

Calling APIs Safely

Making a request and checking the result The requests library is the standard way to call an HTTP API.

6

Making Network Calls Resilient

Setting a timeout on every call By default, requests waits forever for a response.

7

Querying Operational Data with SQLite

Using sqlite3 for quick local data Python includes SQLite, a database that lives in one file or in memory, which makes...

8

Testing Operational Code

Writing unit tests with pytest A unit test checks one function in isolation, with no server or database.

9

Building Command-Line Tools

Turning a script into a tool with argparse A command-line tool takes its input as arguments, answers --help, and can be...

10

Checking Many Things at Once

Why one at a time is too slow Checking a thousand servers one after another takes a thousand times the wait for a...

11

Packaging and Reviewing Operational Code

Isolating dependencies with virtual environments A virtual environment is a private Python installation for one...

12

Building a Budget Reporting Tool in a Hands-on Lab

Preparing the project and the endpoint You will build budget-report, a tool that reads a success rate from an endpoint...

13

Quick Reference and Common Mistakes

Quick reference Keep this table next to you when you write your next operational script.

Career Impact

Roles that use the skills in this module.

  • Site Reliability Engineer

  • DevOps Engineer

  • Platform Engineer

  • Cloud Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Bash is great for short one-liners. Once a task needs retries, error handling, structured output, tests, or has to run unattended every night, Python is easier to make safe and to hand to a teammate.

No. Most SRE checks spend their time waiting on the network, and a thread pool handles that well with far less complexity. Reach for asyncio only when you have a real reason, such as thousands of simultaneous connections.

No. Retry temporary problems such as timeouts, connection errors, 429, and 5xx responses. Fail immediately on client errors such as 401 or 404, because they will fail the same way every time.

Read them from environment variables or a secrets manager at run time, and never write them into code, config files, or log lines. If a secret is ever committed to Git, treat it as leaked and rotate it.