Menu
Explore
Pick a module and start learning specific DevOps tools or concepts at your own pace.
Learn networking for DevOps: IP addresses, subnets, DNS, TCP, HTTPS, firewalls, load balancers and cloud VPCs, then debug real outages step by step.
Learn Linux for DevOps from scratch - file system, permissions, processes, bash scripting, and production server management.
Learn Bash shell scripting for DevOps from zero: variables, loops, functions, error handling, cron, APIs and AWS CLI, with three real projects.
Learn GitLab from first push to production: projects, groups, CI/CD pipelines, variables, runners, and a Docker deploy with a manual approval gate.
Learn Git and GitHub from first commit to team workflow: branching, merging, pull requests, secrets, CI, and recovery tools for when things break.
Learn how code becomes a deployable artifact with Gradle, Poetry, Docker and Makefiles, plus dependency locking and artifact versioning.
Learn Docker from zero: containers, images, Dockerfiles, volumes, networking and Compose, with hands-on labs and projects you can ship.
Learn Kubernetes from zero: pods, deployments, services, ingress, storage, probes, autoscaling, RBAC and Helm, then deploy a real app in a lab.
Learn Azure DevOps from zero: Boards, Repos, YAML pipelines, approvals and service connections, then deploy to Kubernetes with kubectl.
Learn GitHub Actions from zero: workflows, triggers, expressions, matrix builds, secrets, Docker images, OIDC deploys and secure pipelines, with a lab.
A complete, practical reference for DevOps engineers. Covers all 20 core AWS services with architecture, flows, real-world usage, CLI commands, and comparisons.
Learn Azure administration from zero: Entra ID and RBAC, governance, storage, VMs and App Service, networking, monitoring, and backup, mapped to AZ-104.
Learn Google Cloud from zero: projects and IAM, Compute Engine, GKE, Cloud Run, storage, databases, VPCs, and monitoring, mapped to the ACE exam.
Learn Terraform from zero: HCL, providers, state, variables, modules, loops and CI/CD, then build Docker and AWS infrastructure in hands-on labs.
Learn DevOps monitoring and logging from zero: Prometheus, PromQL, Grafana, alerting, Loki, EFK, tracing, Kubernetes and SLOs, with a full hands-on lab.
Learn what DevSecOps means for a DevOps engineer: shift left, the main security scans, who owns what, and a first secure pipeline you build yourself.
Learn Python from absolute zero - variables, data types, loops, functions, files, APIs, and real projects. No prior experience needed.
Learn what changes when Linux runs on AWS: connect with SSM, boot with user data, mount EBS volumes, use IMDSv2, and run apps with systemd.
Learn to automate AWS with Python and boto3: sessions, credentials, paginators, waiters, retries, and safe cleanup scripts with a dry-run default.
Learn CIDR, DNS and Route 53, TCP, TLS, load balancers, NAT, and security groups, then debug AWS connectivity with a six-step chain.
Learn what cloud computing changes: IaaS, PaaS, SaaS, scalability, availability, AWS regions and AZs, and the shared responsibility model.
Learn AWS IAM the safe way: Identity Center, roles, policies, cross-account access, and least privilege, with a lab you can run in CloudShell.
Learn to design and build AWS VPCs: CIDR planning, public and private subnets, route tables, NAT, security groups, endpoints, and flow logs.
Learn Amazon EC2 from instance types and pricing to launch templates, Auto Scaling Groups, and load balancers, with a lab that scales a real app.
Learn AWS storage services: choose and run S3, EBS, EFS, FSx, and transfer tools like DataSync for real production workloads.
Learn AWS databases: RDS, Aurora, DynamoDB, ElastiCache with Valkey, DocumentDB, and Neptune, and how to pick the right one.
Learn Route 53 DNS routing policies, CloudFront CDN, ACM certificates, and WAF - and how to design fast, resilient, global-facing traffic paths.
Learn AWS-native IaC with CloudFormation and CDK, choose between them and Terraform, and run multi-account environments, imports, and CI safely.
Learn to run containers on ECS Fargate and functions on Lambda, expose them with API Gateway, and choose the right compute model for each workload.
Learn to ship to AWS safely: OIDC from CI, images to ECR, ECS and Lambda deploys, blue-green and canary releases, and Terraform pipelines.
Learn Amazon EKS end to end: create clusters with eksctl and Terraform, give pods AWS access, add load balancers and storage, and upgrade safely.
Learn AWS CloudWatch monitoring with metrics, alarms, logs, CloudTrail, Config and EventBridge, and set up alerts that do not cause alert fatigue.
Learn to design resilient AWS architectures and control cloud spend - multi-AZ failure design, RTO/RPO, FinOps, and multi-account strategy.
Learn to move workloads to AWS: choose a migration strategy, rehost with MGN, migrate databases with DMS, and connect on-premises networks.
Learn to design AWS data pipelines that do not bankrupt the company - Kinesis, Glue, Athena, Redshift, and Iceberg explained with real cost numbers.
Learn to choose, deploy, and operate ML infrastructure on AWS - from Rekognition to SageMaker training, deployment modes, drift detection, and Bedrock RAG.
Learn which AWS certification to take first, how to build a portfolio hiring managers trust, and how to prepare for cloud interviews and offers.
Learn the Python AI engineers use daily: types, classes, generators, async calls, retries, config and secrets, tests, and a first FastAPI endpoint.
Learn to load, clean, plot, and query data with NumPy, pandas, and SQL, and to split it without leakage, so every later AI result can be trusted.
Learn just enough math for AI engineering: vectors, matrices, gradients, loss curves, and the statistics behind every evaluation metric you will read.
Learn classical ML for AI engineers: regression, classification, tree models, honest validation, and metrics that tell you when not to use an LLM.
Learn how neural networks learn, why you fine-tune a pre-trained model instead of training from scratch, and how CNNs and Transformers fit in.
Learn how LLMs generate text, how to choose a model on capability, cost, and latency, and how to call and prompt it for reliable structured output.
Learn to turn documents into searchable meaning using embeddings, pgvector, and a real ingestion pipeline with metadata, dedup, and PII redaction.
Learn when fine-tuning beats RAG and prompting, why LoRA is the default, and how to run, evaluate, and check a small LoRA fine-tune for forgetting.
Learn to build a production RAG pipeline: chunking, hybrid search, reranking, parent-child retrieval, citations, and how to debug wrong answers.
Learn to build AI agents that act safely: tool design, explicit state, LangGraph, MCP, retries, idempotency, step limits, and human approval gates.
Learn to prove your AI system works: golden datasets, LLM-as-judge, retrieval and faithfulness metrics, agent evaluation, and release gates in CI.
Learn LLMOps for production AI systems: FastAPI serving, streaming, auth, Docker and Kubernetes basics, tracing, caching, and cost control.
Learn to defend AI products before launch: prompt injection, retrieval authorization, moderation, bias testing, privacy, and a risk-sized checklist.
Explore five AI engineering tracks: advanced agents, multimodal, computer vision, voice, and language AI, then build one starter project in depth.
Learn to prove your AI engineering skills: three portfolio projects for RAG, agents, and a deployed service, plus system design and interview prep.
Master Linux kernel internals for SRE work - process lifecycle, memory management, cgroups, namespaces, CPU scheduler, and production diagnosis using /proc, strace, and eBPF.
Learn to debug production networks like an SRE: TCP states and queues, DNS, TLS, HTTP versions, Kubernetes service paths, and load balancer behaviour.
Learn the security fundamentals every engineer needs: the CIA triad, OWASP Top 10, encryption and TLS, and threat modeling with STRIDE.
Learn to write Python that survives production: safe scripts, resilient API calls, tests, and CLI tools that SRE teams rely on during incidents.
Learn GitOps with Argo CD: sync apps from Git, promote across environments, scale with ApplicationSets, and ship safely with Rollouts canaries.
Learn to run incidents calmly: severity levels, the incident commander role, a six-step debugging method, blameless postmortems, and healthy on-call.
Learn why distributed systems fail in production: CAP, consensus, cascading failures, retry storms, and the patterns that contain them, with labs.
Learn to make failures hurt less: map dependencies, design graceful degradation, and release safely with feature flags, blue-green, and canary releases.
Learn to keep PostgreSQL and Redis reliable: WAL, vacuum, connection pooling, replication lag, failover, point-in-time recovery, and stateful storage.
Learn to break systems on purpose, safely: steady-state hypotheses, blast radius, Chaos Mesh experiments, Game Days, and FMEA to pick what to test.
Learn to load test services with k6, find the real bottleneck, and turn results into a 12 month capacity plan tied to your SLOs.
Learn to measure toil, decide what to remove, and replace it with self-service, GitOps, and safe self-healing automation, proven with real data.
Learn senior SRE design work: reliability reviews, FMEA, production readiness reviews, RTO and RPO, and multi-region disaster recovery you can prove.
Learn to scale reliability beyond your own services: golden paths with safe defaults, catalog ownership, platform SLOs, and adoption you can measure.
Learn to lead reliability without authority: run blameless reviews, grow on-call engineers, shape designs early, and win support for reliability work.
Learn to answer SRE interview questions with confidence: SLO design, incident walkthroughs, reliability system design, and a job-ready checklist.
Learn to keep secrets out of Git with push protection, pre-commit hooks, and Gitleaks, protect main with rulesets, sign commits, and respond to leaks.
Learn to harden CI/CD: OIDC instead of stored keys, scoped tokens, SHA pinned actions, Vault secrets, Jenkins hardening, signing, and audit logs.
Learn to catch code and dependency flaws early with SAST and SCA, prioritise fixes with CVSS and EPSS, and replace static secrets with Vault.
Learn to run DAST with OWASP ZAP, generate SBOMs with Syft, scan with Grype, and sign and verify images with Cosign and SLSA provenance.
Learn to secure Kubernetes with least-privilege RBAC, Pod Security Standards, network policies, encrypted secrets, and runtime alerts on a local cluster.
Learn to secure AWS workloads end to end: threat detection with GuardDuty, encryption with KMS, secrets management, WAF, and incident response.
Learn zero trust identity in practice: SSO and MFA for people, federation for workloads, just-in-time access, short-lived certificates, and OAuth flows.
Learn to detect attacks in running containers with Falco, map detections to MITRE ATT&CK, test them, and route alerts to a safe, automated response.
Learn how SOC 2, ISO 27001, and PCI DSS map to your pipelines, automate audit evidence, and run security incidents from detection to postmortem.
Learn advanced DevSecOps: test web and API attacks, hunt threats in your logs, automate security response, and secure AI-powered applications.
Learn to prove your DevSecOps skills: build one complete secure pipeline, document it like a professional, and prepare for interviews at every level.
Learn how the apps you run actually work: HTTP and REST, reverse proxies, databases, config, and health checks, seen from the platform side.
Learn to build hardened Docker images, keep secrets out of layers, scan with Trivy, and run containers with least privilege, read-only files, and seccomp.
Learn to set an AWS security baseline: Identity Center, guardrails with SCPs, CloudTrail, Config, Access Analyzer, and Security Hub across accounts.
Learn to secure Terraform: keep secrets out of code and state, pin providers and modules, scan with Checkov and Trivy, and enforce rules with Conftest.
Learn to run Kubernetes for many teams: tenant namespaces, quotas, Kyverno guardrails, Crossplane self-service, and cluster upgrades done safely.
Learn when a service mesh is worth it, and how Istio adds mTLS, traffic shifting, and resilience without code changes, compared with Linkerd and ambient.
Learn to build an internal developer platform as a product: a Backstage catalog, golden path templates, self-service infrastructure, and DORA metrics.
Learn FinOps for Kubernetes: see costs by team with OpenCost, rightsize workloads, use Spot safely with Karpenter, and build showback teams trust.
Learn to prove your platform skills: finish the acme-shop paved road project, design internal platforms on a whiteboard, and prepare for interviews.
Learn to write Python pipeline scripts that run unattended: pandas, files, APIs, logging, tests, and idempotent writes, the glue of every data stack.
Learn the SQL data engineers write daily: joins, aggregation, window functions, CTEs, upserts, query plans, and validation checks on real tables.
Learn to design tables people can trust: OLTP versus OLAP, normalization, grain, star schemas, slowly changing dimensions, and double counting.
Learn how warehouses read less data to answer faster: columnar storage, parallel processing, pruning, partitioning, and clustering, proven with DuckDB.
Learn the AWS services data engineers use daily: an S3 data lake, the Glue catalog and jobs, Athena, Redshift, and least-privilege IAM for pipelines.
Learn to build Airflow 3 pipelines that survive real failures: DAGs, retries, idempotent loads, backfills, watermarks, CDC concepts, and quality gates.
Learn Docker for data work: build pipeline images, persist data with volumes, connect services, and run Postgres, Kafka, and Airflow with Compose.
Learn to build tested, documented SQL transformations with dbt: staging to marts, materializations, incremental models, tests, and lineage.
Learn PySpark for data too big for one machine: DataFrames, partitions, shuffles, joins, the Spark UI, and the fixes that make slow jobs fast.
Learn how cloud warehouses and Delta Lake power the modern stack: separate compute, time travel, ACID tables, MERGE, and the Medallion layers.
Learn Flink SQL to turn Kafka streams into live results: windows, event time, watermarks, state, and checkpoints, at a working level.
Learn Kafka for data pipelines: topics, partitions, consumer groups, Kafka Connect, CDC with Debezium, Schema Registry, and how Kinesis compares.
Learn practical data governance: classify data, catalog and lineage, role and row access, masking, retention, and erasure under GDPR and DPDP.
Learn to stop bad data before anyone trusts it: Great Expectations suites, checkpoints, severity rules, quarantine, data contracts, and Airflow gates.
Learn to test data pipelines with pytest and dbt, block bad pull requests in CI, deploy Airflow DAGs with GitHub Actions, and roll back safely.
Explore the main data engineering directions: analytics engineering, data platforms, ML and AI data infrastructure, and streaming, then choose yours.
Learn to build three data engineering portfolio projects, document them like a pro, and answer data engineer interview questions with confidence.
Learn to instrument services with OpenTelemetry, run the Collector, and follow one slow request across services to find the real cause of an incident.
Learn to define SLIs and SLOs, track error budgets, and write multi-window burn-rate alerts that page only when users are really affected.
Learn to move ops data where models can use it: stream alerts and metrics through Kafka, export Prometheus history, and check data quality first.
Learn to catch ops anomalies early: rolling and seasonal baselines first, then Isolation Forest, with precision and recall proving fewer false pages.
Learn to forecast ops metrics with Prophet and ARIMA: predict capacity exhaustion, alert before thresholds break, and flag anomalies with bands.
Learn to turn alert storms into single incidents with Alertmanager grouping, inhibition, and silences, then measure MTTD and MTTR to prove it works.
Learn a structured way to find root cause during incidents: timeline, blast radius, signals, hypotheses, and verified fixes, then a useful postmortem.
Learn what commercial AIOps platforms automate, when the open-source stack is enough, and how to judge an AI investigation by speed and accuracy.
Learn to prompt LLMs for ops work: classify alerts with few-shot examples, reason from symptoms to hypotheses, and return JSON automation can trust.
Use seven ready prompt templates for RCA, runbooks, stakeholder updates, and postmortems, filled from real incident data and chained through an incident.
Learn to choose models for ops work by quality, latency, and cost, route cheap and strong models, and run local models with Ollama for sensitive logs.
Learn how ops agents work: the reason-act-observe loop, tool definitions, workflows versus agents, and building an investigator with LangGraph.
Learn to ground ops agents in your own runbooks and postmortems: chunk, embed, retrieve with citations, and know when live data beats retrieval.
Learn to build an MCP server that gives AI agents safe, read-only access to Kubernetes, Prometheus, and runbooks, with least-privilege RBAC.
Learn to make ops agents safe: allow-list parsed commands, gate risky actions with approval, cap steps and cost, audit every call, and resist injection.
Learn to fix known failures automatically: Kubernetes self-healing first, then idempotent Ansible runbooks triggered by alerts, with approval and audit.
Learn to prove AIOps skills with three portfolio projects: anomaly detection, a RAG ops assistant, and a safe investigation agent, plus interview prep.
Automate file transfer from Amazon S3 to FSx for OpenZFS so that every file uploaded to a specific S3 path automatically appears on FSx - no manual commands needed after setup.