Production Reliability Engineering (SRE) Course | Ingress Ac
Create account & apply
Site Reliability Engineer DEVOPS & LINUX ENGINEERING · ADVANCED LEVEL

Site Reliability Engineering (SRE) Bootcamp

An intensive 2-month, hands-on SRE program that teaches you how to measure, monitor, and defend production reliability in modern cloud-native and Kubernetes environments.

Advanced8 weeks48 hoursOnline
Your selected course will be carried into Ingress Portal.
COURSE HIGHLIGHTS
πŸ“Š

Metrics-Driven Reliability

Learn to define SLIs, SLOs, and error budgets that turn reliability into a measurable, actionable engineering practice.

🚨

Real Incident Simulations

Practice alerting, RCA, and blameless postmortems through realistic production incident scenarios, not just theory.

☸️

Kubernetes-Native Reliability

Apply advanced Kubernetes patterns like PDBs, affinity rules, and progressive rollouts to build resilient production systems.

Not sure you are ready? Take a skill assessment
YOUR LEVEL
BeginnerIntermediateAdvancedExpert
WHAT YOU WILL BE ABLE TO DO

The skills you will have by the end of this course.

  • Define and calculate SLIs, SLOs, SLAs, and error budgets for production services
  • Build and interpret observability stacks using Prometheus, Grafana, logs, and traces
  • Design effective alerting strategies and manage the full incident lifecycle
  • Apply Kubernetes reliability patterns including probes, autoscaling, PDBs, and deployment strategies
  • Implement GitOps-based, safe software delivery with rollback and progressive delivery techniques
  • Plan disaster recovery strategies and execute chaos engineering and load testing exercises
  • Lead root cause analysis and produce blameless postmortems and runbooks
  • Leverage AI-assisted tools for troubleshooting, RCA, and incident response
CONDENSED SYLLABUS

See the structure without reading a textbook.

Modules stay collapsed for quick scanning. Open any module to inspect its topics.

01Introduction to Site Reliability Engineering6 lessons+

What is SRE? Evolution from Operations to DevOps to SRE

DevOps vs SRE and Reliability Engineering Principles

Production Mindset and Service Lifecycle

Shared Responsibility Model, Toil, and Automation

SRE Roles and Responsibilities

Lab: Calculate Availability and Identify Toil

02Measuring Service Reliability6 lessons+

Why Reliability Needs Metrics: Availability vs Reliability

Service Level Indicators (SLI) and Objectives (SLO)

Service Level Agreements (SLA) and Error Budgets

MTTD, MTTR, and MTBF

DORA Metrics for Delivery Performance

Lab: Define SLOs and Calculate Error Budgets

03Monitoring & Observability6 lessons+

Monitoring vs Observability and the Three Pillars

Metrics, Logs, and Traces Fundamentals

Golden Signals, RED Method, and USE Method

Metrics Collection with Prometheus

Visualizing System Health with Grafana and Distributed Tracing

Lab: Analyze Dashboards and Detect Bottlenecks

04Alerting & Incident Management6 lessons+

Alerting Principles and Prometheus Alertmanager

Alert Fatigue, Severity, and Prioritization

Incident Lifecycle and Response Process

On-call Best Practices and Blameless Postmortems

Root Cause Analysis and Writing Effective Runbooks

Lab: Simulate Incidents, Perform RCA, and Write a Postmortem

05Kubernetes Reliability Engineering6 lessons+

Self-Healing, Liveness, Readiness, and Startup Probes

Resource Requests, Limits, and Horizontal/Vertical Autoscaling

Pod Disruption Budgets and Scheduling for Reliability

Node Affinity, Anti-Affinity, and Topology Spread Constraints

Reliable Deployment Strategies: Rolling, Blue-Green, Canary

Lab: Troubleshoot Pods, Simulate Node Failure, and Test HPA

06Reliable Software Delivery6 lessons+

CI/CD Reliability and GitOps Principles

Progressive Delivery and Safe Deployment Practices

Rollback vs Roll-forward and Release Management

Configuration and Secret Management

Change Management, Deployment Verification, and Supply Chain Security

Lab: GitOps Sync and Safe Production Release Simulation

07Production Resilience & Disaster Recovery6 lessons+

Disaster Recovery Fundamentals and High Availability Architecture

Backup & Restore, RTO, and RPO

Single Point of Failure and Capacity Planning Review

Load Testing Concepts and Chaos Engineering

Recovery Strategies: Active-Active, Active-Passive, Warm Standby, Pilot Light

Lab: Restore from Backup and Run Chaos Testing Scenarios

08Modern SRE Operations & Production Scenarios6 lessons+

End-to-End Troubleshooting Methodology

Common Production Failures and Case Studies

Cost vs Reliability Trade-offs

AI for SRE: AIOps, AI-assisted Troubleshooting and RCA

SRE Career Roadmap and Continuous Improvement

Capstone Lab: Full Production Incident Simulation

UPCOMING GROUPS

Choose the cohort you can actually attend.

Only current, open groups are shown.

STARTS

To be announced

Online
Schedule
To be announced
Format
Online
Duration
8 weeks · 48 hours
Language
Confirm with advisor
Join the next cohort waitlist
CONTEXTUAL PROOF
“The program helped me connect individual skills into the way real teams design, build and deliver software.”

See real graduate stories from the Ingress community.

Explore graduate results
APPLICATION THROUGH INGRESS PORTAL

Your course stays selected while you create your account.

We use one Portal account for applications, assessments and future learning progress. You will not need to email your details or select the training again.

Questions first? Talk to an advisor
  1. 01

    Create or sign in to your Portal accountYour contact details stay connected to one student profile.

  2. 02

    Confirm your application detailsSite Reliability Engineering (SRE) Bootcamp is preselected.

  3. 03

    Submit your applicationThe admissions team receives it immediately and can follow up from the Portal.

SELECTED TRAININGSite Reliability Engineering (SRE) BootcampOnline

Continue in Ingress Portal Already registered? The Portal will let you sign in instead.