Metrics-Driven Reliability
Learn to define SLIs, SLOs, and error budgets that turn reliability into a measurable, actionable engineering practice.
Hesab yarat və müraciət et
An intensive 2-month, hands-on SRE program that teaches you how to measure, monitor, and defend production reliability in modern cloud-native and Kubernetes environments.
Learn to define SLIs, SLOs, and error budgets that turn reliability into a measurable, actionable engineering practice.
Practice alerting, RCA, and blameless postmortems through realistic production incident scenarios, not just theory.
Apply advanced Kubernetes patterns like PDBs, affinity rules, and progressive rollouts to build resilient production systems.
Modullar sürətli baxış üçün yığcam qalır. Mövzularına baxmaq üçün istənilən modulu açın.
What is SRE? Evolution from Operations to DevOps to SRE
DevOps vs SRE and Reliability Engineering Principles
Production Mindset and Service Lifecycle
Shared Responsibility Model, Toil, and Automation
SRE Roles and Responsibilities
Lab: Calculate Availability and Identify Toil
Why Reliability Needs Metrics: Availability vs Reliability
Service Level Indicators (SLI) and Objectives (SLO)
Service Level Agreements (SLA) and Error Budgets
MTTD, MTTR, and MTBF
DORA Metrics for Delivery Performance
Lab: Define SLOs and Calculate Error Budgets
Monitoring vs Observability and the Three Pillars
Metrics, Logs, and Traces Fundamentals
Golden Signals, RED Method, and USE Method
Metrics Collection with Prometheus
Visualizing System Health with Grafana and Distributed Tracing
Lab: Analyze Dashboards and Detect Bottlenecks
Alerting Principles and Prometheus Alertmanager
Alert Fatigue, Severity, and Prioritization
Incident Lifecycle and Response Process
On-call Best Practices and Blameless Postmortems
Root Cause Analysis and Writing Effective Runbooks
Lab: Simulate Incidents, Perform RCA, and Write a Postmortem
Self-Healing, Liveness, Readiness, and Startup Probes
Resource Requests, Limits, and Horizontal/Vertical Autoscaling
Pod Disruption Budgets and Scheduling for Reliability
Node Affinity, Anti-Affinity, and Topology Spread Constraints
Reliable Deployment Strategies: Rolling, Blue-Green, Canary
Lab: Troubleshoot Pods, Simulate Node Failure, and Test HPA
CI/CD Reliability and GitOps Principles
Progressive Delivery and Safe Deployment Practices
Rollback vs Roll-forward and Release Management
Configuration and Secret Management
Change Management, Deployment Verification, and Supply Chain Security
Lab: GitOps Sync and Safe Production Release Simulation
Disaster Recovery Fundamentals and High Availability Architecture
Backup & Restore, RTO, and RPO
Single Point of Failure and Capacity Planning Review
Load Testing Concepts and Chaos Engineering
Recovery Strategies: Active-Active, Active-Passive, Warm Standby, Pilot Light
Lab: Restore from Backup and Run Chaos Testing Scenarios
End-to-End Troubleshooting Methodology
Common Production Failures and Case Studies
Cost vs Reliability Trade-offs
AI for SRE: AIOps, AI-assisted Troubleshooting and RCA
SRE Career Roadmap and Continuous Improvement
Capstone Lab: Full Production Incident Simulation
Yalnız cari, açıq qruplar göstərilir.
“The program helped me connect individual skills into the way real teams design, build and deliver software.”
Ingress icmasından real məzun hekayələrinə baxın.
Məzun nəticələrini kəşf etWe use one Portal account for applications, assessments and future learning progress. You will not need to email your details or select the training again.
Əvvəlcə sualınız var? Məsləhətçi ilə danışınPortal hesabınızı yaradın və ya daxil olunƏlaqə məlumatlarınız vahid tələbə profilinə bağlı qalır.
Müraciət məlumatlarınızı təsdiqləyinSite Reliability Engineering (SRE) Bootcamp əvvəlcədən seçilib.
Müraciətinizi göndərinQəbul komandası müraciəti dərhal alır və Portal vasitəsilə əlaqə saxlaya bilər.
SEÇİLMİŞ TƏLİMSite Reliability Engineering (SRE) BootcampOnlayn
Ingress Portalda davam et Artıq qeydiyyatdan keçmisiniz? Portal daxil olmağınıza imkan verəcək.