"How reliable should our service be?" Without a shared answer, product teams push for speed, operations teams push for stability, and every incident becomes an argument.
SLOs, SLIs and error budgets, popularised by Google's Site Reliability Engineering practices, turn that argument into a number everyone agrees on.
What is an SLI?
A service level indicator is a quantitative measure of one aspect of service quality from the user's perspective. It is usually expressed as a ratio of good events to total events.
Common SLIs:
- Availability: successful requests ÷ total requests
- Latency: requests faster than 300 ms ÷ total requests
- Freshness: data updated within 10 minutes ÷ total checks
- Correctness: correct results ÷ total results
What is an SLO?
A service level objective is a target value for an SLI over a defined window. Examples:
- 99.9% of checkout requests succeed over a rolling 30 days
- 95% of search requests complete in under 300 ms over 28 days
What is the difference between SLO, SLI and SLA?
| Term | What it is | Audience | Example |
|---|---|---|---|
| SLI | A measurement | Engineering | % of successful requests |
| SLO | An internal target | Engineering and product | 99.9% success over 30 days |
| SLA | A contractual promise with consequences | Customers | 99.5% uptime or service credits |
SLOs should be stricter than SLAs, so you have room to react before a contract is breached.
SLOs should be stricter than SLAs, so you have room to react before a contract is breached.
What is an error budget?
The error budget is 100% minus the SLO. A 99.9% availability SLO over 30 days allows 0.1% of requests to fail. In time terms, that is roughly 43 minutes of full downtime per 30 days.
Error budget burn-down chart
Illustration in progress
| SLO | Error budget | Approx. downtime per 30 days |
|---|---|---|
| 99% | 1% | 7.2 hours |
| 99.5% | 0.5% | 3.6 hours |
| 99.9% | 0.1% | 43 minutes |
| 99.95% | 0.05% | 22 minutes |
| 99.99% | 0.01% | 4.3 minutes |
When budget remains, teams can ship faster and take risks. When the budget is exhausted, the agreed policy might be to pause risky releases and focus on reliability work.
How do you choose good SLIs?
- Start from user journeys, such as login, search and checkout, not from infrastructure.
- Measure as close to the user as possible, for example at the load balancer or with real user monitoring.
- Keep it to two or three SLIs per service.
- Use ratios of good to total events so they are easy to reason about.
How do you set realistic SLO targets?
- Look at historical performance and set the target slightly below what you already achieve.
- Ask what level of reliability users and the business actually need.
- Remember dependencies: your service cannot be more reliable than the critical services it depends on.
- Review targets every quarter.
What is burn-rate alerting?
Burn rate is how fast you are consuming the error budget relative to the sustainable pace. A burn rate of 1 uses exactly the whole budget over the window; a burn rate of 10 would exhaust it in a tenth of the time.
Multi-window, multi-burn-rate alerts, described in Google's SRE Workbook, page quickly for fast burns (for example, a high burn rate over 1 hour confirmed by a 5-minute window) and open tickets for slow burns. This approach catches real problems while avoiding the noise of static thresholds, and it is one of the most effective tools against alert fatigue.
How Crozaint approaches SLOs
SLO design is a core part of the Operationalize phase (weeks 5–6) in Crozaint's 8-week observability engagement. We work with your product and engineering teams to choose SLIs for critical journeys, set targets, define error budget policies and implement burn-rate alerting on Grafana Cloud.
After handoff, continuous improvement cycles can follow ITIL 4 or SRE frameworks, depending on how your organisation operates. Crozaint's managed services clients have seen uptime of 99.98% in recent reporting periods.
Common mistakes to avoid
- Setting 100% or 99.999% targets without business need
- Using CPU or memory as SLIs
- Defining SLOs but never agreeing an error budget policy
- Too many SLOs per service
- Never reviewing targets as the product changes
Conclusion
SLOs give engineering and product a shared definition of "reliable enough." Measure what users feel, set honest targets, agree a budget policy and alert on burn rate.
Want SLOs your teams actually use? Book a 30-minute discovery call with Crozaint.
