← 📈 Observability & Reliability
Advanced
Multi-window burn-rate alerts: budget arithmetic and recovery signals
Derive burn-rate thresholds from an error budget, combine long and short windows correctly and handle low traffic or missing telemetry without misleading pages.
View Premium accesscourse lessons, concepts and answers · see current terms
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS
Observability, SLOs & ReliabilityDefine SLI, SLO and SLA, compute the monthly error budget for 99.95 percent, and tell me what changes when it is spent.→Observability, SLOs & ReliabilityOn-call is drowning in alerts and starting to ignore them. Redesign the alerting.→Observability, SLOs & ReliabilityWrite the PromQL behind our availability SLO and tell me where these queries usually go wrong.→Observability, SLOs & ReliabilityThe error budget ran out. Describe exactly what happens next and what makes the policy stick.→Observability, SLOs & ReliabilityHow does Prometheus collect metrics? Explain the pull model, exporters and when you need a pushgateway.→Observability, SLOs & ReliabilityDesign a chaos engineering programme. What do you inject first, and how do you avoid causing the outage you were preventing?→