0%

Preparing the page

SLOs, SLIs and Error Budgets Explained (With Practical Examples)

SLOs, SLIs and error budgets explained with examples. Learn how to choose SLIs, set realistic SLO targets and use burn-rate alerts to balance speed and reliability.

Girish

· 5 min read

Share
Placeholder illustration

"How reliable should our service be?" Without a shared answer, product teams push for speed, operations teams push for stability, and every incident becomes an argument.

SLOs, SLIs and error budgets, popularised by Google's Site Reliability Engineering practices, turn that argument into a number everyone agrees on.

What is an SLI?

A service level indicator is a quantitative measure of one aspect of service quality from the user's perspective. It is usually expressed as a ratio of good events to total events.

Common SLIs:

  • Availability: successful requests ÷ total requests
  • Latency: requests faster than 300 ms ÷ total requests
  • Freshness: data updated within 10 minutes ÷ total checks
  • Correctness: correct results ÷ total results

What is an SLO?

A service level objective is a target value for an SLI over a defined window. Examples:

  • 99.9% of checkout requests succeed over a rolling 30 days
  • 95% of search requests complete in under 300 ms over 28 days

What is the difference between SLO, SLI and SLA?

TermWhat it isAudienceExample
SLIA measurementEngineering% of successful requests
SLOAn internal targetEngineering and product99.9% success over 30 days
SLAA contractual promise with consequencesCustomers99.5% uptime or service credits

SLOs should be stricter than SLAs, so you have room to react before a contract is breached.

SLOs should be stricter than SLAs, so you have room to react before a contract is breached.

What is an error budget?

The error budget is 100% minus the SLO. A 99.9% availability SLO over 30 days allows 0.1% of requests to fail. In time terms, that is roughly 43 minutes of full downtime per 30 days.

Error budget burn-down over 30 days
SLOError budgetApprox. downtime per 30 days
99%1%7.2 hours
99.5%0.5%3.6 hours
99.9%0.1%43 minutes
99.95%0.05%22 minutes
99.99%0.01%4.3 minutes

When budget remains, teams can ship faster and take risks. When the budget is exhausted, the agreed policy might be to pause risky releases and focus on reliability work.

How do you choose good SLIs?

  1. Start from user journeys, such as login, search and checkout, not from infrastructure.
  2. Measure as close to the user as possible, for example at the load balancer or with real user monitoring.
  3. Keep it to two or three SLIs per service.
  4. Use ratios of good to total events so they are easy to reason about.

How do you set realistic SLO targets?

  • Look at historical performance and set the target slightly below what you already achieve.
  • Ask what level of reliability users and the business actually need.
  • Remember dependencies: your service cannot be more reliable than the critical services it depends on.
  • Review targets every quarter.

What is burn-rate alerting?

Burn rate is how fast you are consuming the error budget relative to the sustainable pace. A burn rate of 1 uses exactly the whole budget over the window; a burn rate of 10 would exhaust it in a tenth of the time.

Multi-window, multi-burn-rate alerts, described in Google's SRE Workbook, page quickly for fast burns (for example, a high burn rate over 1 hour confirmed by a 5-minute window) and open tickets for slow burns. This approach catches real problems while avoiding the noise of static thresholds, and it is one of the most effective tools against alert fatigue.

How Crozaint approaches SLOs

SLO design is a core part of the Operationalize phase (weeks 5–6) in Crozaint's 8-week observability engagement. We work with your product and engineering teams to choose SLIs for critical journeys, set targets, define error budget policies and implement burn-rate alerting on Grafana Cloud.

After handoff, continuous improvement cycles can follow ITIL 4 or SRE frameworks, depending on how your organisation operates. Crozaint's managed services clients have seen uptime of 99.98% in recent reporting periods.

Common mistakes to avoid

  • Setting 100% or 99.999% targets without business need
  • Using CPU or memory as SLIs
  • Defining SLOs but never agreeing an error budget policy
  • Too many SLOs per service
  • Never reviewing targets as the product changes

Conclusion

SLOs give engineering and product a shared definition of "reliable enough." Measure what users feel, set honest targets, agree a budget policy and alert on burn rate.

Want SLOs your teams actually use? Book a 30-minute discovery call with Crozaint.

Frequently Asked Questions

What is the difference between an SLO and an SLA?

An SLO is an internal reliability target, such as 99.9% successful requests over 30 days. An SLA is a contractual commitment to customers, usually with financial penalties if missed. SLOs are normally set stricter than SLAs so teams can act before the contract is breached.

What is a good SLO target?

There is no universal target. Many user-facing services use 99.5% to 99.9% availability. The right target reflects user expectations, business impact, historical performance and dependency reliability. Start slightly below current performance and adjust quarterly as you learn.

How do you calculate an error budget?

Subtract the SLO from 100%. A 99.9% SLO gives a 0.1% error budget. Over 30 days, that equals roughly 43 minutes of full downtime, or 1 failed request per 1,000. The budget is consumed by failures and replenishes as the rolling window moves.

What happens when the error budget runs out?

That depends on your error budget policy. Common responses include freezing risky feature releases, prioritising reliability work, requiring extra review for changes and running a postmortem. The policy should be agreed in advance by engineering and product leadership.

Do SLOs replace monitoring alerts?

SLO burn-rate alerts replace many symptom and threshold alerts for paging, because they focus on user impact. You still need some infrastructure monitoring for capacity planning and specific failure modes, but those alerts often become tickets or dashboards rather than pages.

Written by

Girish

Crozaint · 15 articles

Full profile coming soon.

JosephReviewed for technical accuracy by Joseph, Cloud Consulting.

After the reading

Reading About Observability Is the Easy Part.Doing It in Your Estate Is Ours.

Thirty minutes with the people who wrote this. We look at your setup, say what we would fix first and leave you with a plan, whether or not you go further with us.

  • A look at your estate, not a demo
  • What we would fix first, and why
  • A plan you keep, whether or not you hire us
Joseph

Talk to Joseph

Cloud Consulting

Thirty minutes on your estate. Joseph looks at what you have and tells you what we would do first.

Book 30 Minutes

No deck, no pitch, no commitment.