SLO: [service name]
Service Level Objective – the written contract between the service
and its users' patience. Adapted from Google SRE practice.
| Field | Value |
|---|
| Service | [name] |
| Owner | @name |
| Approvers | @name, @name |
| Window | Rolling 30 days |
| Effective | YYYY-MM-DD |
| Status | Draft / Active / Deprecated |
SLIs (indicators)
What is measured, exactly. "Not returning 5xx" is not availability;
define the narrow, measurable event.
| SLI | Definition | Data source |
|---|
| Availability | 2xx + 3xx responses / total responses | Load balancer logs |
| Latency | p95 request duration < 300 ms | APM traces |
| Freshness | data lag < 60 s | Pipeline metrics |
SLOs (targets)
| SLI | Target | Window | Error budget |
|---|
| Availability | 99.9% | 30d rolling | ~43 min downtime / 30d |
| Latency | p95 < 300 ms | 30d rolling | 5% of requests may exceed |
A 100% target is neither realistic nor desirable – it removes the
headroom that lets engineering ship.
Error budget policy
| Budget remaining | Action |
|---|
| > 50% | Normal feature velocity |
| 25-50% | Reliability work takes priority |
| < 25% | Feature freeze; reliability work only |
| Exhausted | Releases stopped; escalate to @name |
Alerting (burn rate)
| Window | Burn rate | Severity |
|---|
| 1h / 5m | > 14x | Page on-call |
| 6h / 30m | > 6x | Page on-call |
| 3d / 6h | > 1x | File ticket |
Reporting & review
- Dashboard: [link]
- Review cadence: monthly / quarterly
- Escalation path: @name -> @name
Revision history
| Version | Date | Change | Author |
|---|
| 1.0 | YYYY-MM-DD | Initial | @name |