Metrics-Driven Reliability
Learn to define SLIs, SLOs, and error budgets that turn reliability into a measurable, actionable engineering practice.
Создать аккаунт и подать заявку
An intensive 2-month, hands-on SRE program that teaches you how to measure, monitor, and defend production reliability in modern cloud-native and Kubernetes environments.
Learn to define SLIs, SLOs, and error budgets that turn reliability into a measurable, actionable engineering practice.
Practice alerting, RCA, and blameless postmortems through realistic production incident scenarios, not just theory.
Apply advanced Kubernetes patterns like PDBs, affinity rules, and progressive rollouts to build resilient production systems.
Модули свёрнуты для быстрого просмотра. Откройте любой модуль, чтобы увидеть его темы.
What is SRE? Evolution from Operations to DevOps to SRE
DevOps vs SRE and Reliability Engineering Principles
Production Mindset and Service Lifecycle
Shared Responsibility Model, Toil, and Automation
SRE Roles and Responsibilities
Lab: Calculate Availability and Identify Toil
Why Reliability Needs Metrics: Availability vs Reliability
Service Level Indicators (SLI) and Objectives (SLO)
Service Level Agreements (SLA) and Error Budgets
MTTD, MTTR, and MTBF
DORA Metrics for Delivery Performance
Lab: Define SLOs and Calculate Error Budgets
Monitoring vs Observability and the Three Pillars
Metrics, Logs, and Traces Fundamentals
Golden Signals, RED Method, and USE Method
Metrics Collection with Prometheus
Visualizing System Health with Grafana and Distributed Tracing
Lab: Analyze Dashboards and Detect Bottlenecks
Alerting Principles and Prometheus Alertmanager
Alert Fatigue, Severity, and Prioritization
Incident Lifecycle and Response Process
On-call Best Practices and Blameless Postmortems
Root Cause Analysis and Writing Effective Runbooks
Lab: Simulate Incidents, Perform RCA, and Write a Postmortem
Self-Healing, Liveness, Readiness, and Startup Probes
Resource Requests, Limits, and Horizontal/Vertical Autoscaling
Pod Disruption Budgets and Scheduling for Reliability
Node Affinity, Anti-Affinity, and Topology Spread Constraints
Reliable Deployment Strategies: Rolling, Blue-Green, Canary
Lab: Troubleshoot Pods, Simulate Node Failure, and Test HPA
CI/CD Reliability and GitOps Principles
Progressive Delivery and Safe Deployment Practices
Rollback vs Roll-forward and Release Management
Configuration and Secret Management
Change Management, Deployment Verification, and Supply Chain Security
Lab: GitOps Sync and Safe Production Release Simulation
Disaster Recovery Fundamentals and High Availability Architecture
Backup & Restore, RTO, and RPO
Single Point of Failure and Capacity Planning Review
Load Testing Concepts and Chaos Engineering
Recovery Strategies: Active-Active, Active-Passive, Warm Standby, Pilot Light
Lab: Restore from Backup and Run Chaos Testing Scenarios
End-to-End Troubleshooting Methodology
Common Production Failures and Case Studies
Cost vs Reliability Trade-offs
AI for SRE: AIOps, AI-assisted Troubleshooting and RCA
SRE Career Roadmap and Continuous Improvement
Capstone Lab: Full Production Incident Simulation
Показаны только текущие открытые группы.
“Программа помогла мне связать отдельные навыки с тем, как реальные команды проектируют, создают и поставляют ПО.”
Смотрите реальные истории выпускников из сообщества Ingress.
Изучить результаты выпускниковМы используем один аккаунт Portal для заявок, оценок и будущего прогресса обучения. Вам не нужно отправлять данные по почте или снова выбирать курс.
Сначала вопросы? Поговорите с консультантомСоздайте аккаунт Portal или войдитеВаши контактные данные привязаны к одному профилю студента.
Подтвердите данные заявкиSite Reliability Engineering (SRE) Bootcamp уже выбран.
Отправьте заявкуПриёмная команда получает её сразу и может связаться с вами через Portal.
ВЫБРАННЫЙ КУРСSite Reliability Engineering (SRE) BootcampОнлайн
Продолжить в Ingress Portal Уже зарегистрированы? Portal позволит вам войти.