Runbook: Service name
On-call runbook – alerts, procedures, escalation
Mode d'emploi: Anything you'll be paged for at 3 AM. Write for the version of you that just woke up.
Aperçu
Runbook: Service name
| Field | Value |
|---|---|
| Service | service-name |
| Owner | Team / @person |
| On-call | Rotation link |
| Dashboard | Monitoring link |
| Repo | Source link |
Service overview
What the service does, what depends on it, what it depends on.
flowchart LR
LB[Load balancer] --> S[service-name]
S --> DB[(postgres)]
S --> Cache[(redis)]
S --> Q[queue]
Alerts and responses
Alert: high error rate
- Threshold: >2% 5xx for 5 min
- Check:
kubectl logs -l app=service-name --tail=100 - Common cause: downstream DB connection exhaustion
- Fix: scale replicas / restart pooler
Alert: latency p99 high
- Threshold: p99 > 800ms for 10 min
- Check: dashboard -> slow queries panel
- Fix: see "Rollback" below if tied to recent deploy
Procedures
Deploy
git tag v1.2.3 && git push --tags
# CI builds and pushes the image
ssh prod 'docker compose pull && docker compose up -d'
Rollback
# in .env on the host
APP_VERSION=<previous-tag>
docker compose up -d
Health check
curl -fsS https://service.example.com/api/health
Escalation
- On-call engineer (PagerDuty)
- Team lead @lead
- SRE manager @manager