A service level objective (SLO) is a specific, measurable target for how well a service should perform, for example “99.9% of checkout requests succeed over a rolling 30 days” or “95% of pages load in under one second”. It sits between the raw measurement, called a service level indicator (SLI), and the contractual promise, called a service level agreement (SLA). SLOs are mostly internal targets that tell a team, and the business, when reliability is good enough and when it needs attention.
At a glance
- An SLO is a target for a measurement (the SLI) over a time window, such as success rate or response time over 30 days.
- It is usually an internal objective; an SLA is typically a contract with remedies.
- Good SLOs measure what users experience, not just whether a server is powered on.
- Internal SLOs are commonly set tighter than external SLAs to give early warning.
- The gap between the target and 100% is often treated as an error budget for changes and risk.
What problem it solves
Many organizations judge reliability by an uptime percentage in a contract, then find that the service meets it while users still complain. A link can be “up” while dropping packets; an application can respond while every page takes ten seconds. SLAs also tend to be loose, because providers tend to sign only what they are confident they can meet.
SLOs fix this by agreeing, in advance, what “working” means for a service from the user’s point of view and how much failure is acceptable. That gives engineering, operations and business leaders one shared number to manage against. It replaces arguments about whether last week was “bad” with a measured answer, and it tells you when to stop adding features and invest in stability instead.
How it works
Choose the indicators. A team first picks service level indicators: measurements that reflect what users care about. Common SLIs include the share of requests that succeed, the share answered within a latency threshold, data freshness for pipelines, or call quality for voice. Good SLIs are measured as close to the user as practical, often with tools for application performance monitoring and observability.
Set the target and window. The SLO adds a target and a period: 99.9% of requests successful over 28 or 30 days, or 99% of API calls under 300 milliseconds. Targets reflect what users need and what the business will pay for, not the highest number the system has ever achieved.
Track the error budget. Whatever the SLO leaves on the table is the error budget. A 99.9% monthly target allows roughly 43 minutes of failure in a 30-day month. Teams following site reliability engineering (SRE) practices watch how fast that budget is being spent, alert when it burns too quickly, and agree in advance what happens when it runs out, such as pausing risky releases.
Review and adjust. SLOs are revisited as services, users and costs change. A target nobody ever misses may be too loose; one missed every month may be unrealistic or point to real underinvestment.
When it matters for buyers
- When negotiating an SLA. Knowing your own SLOs tells you what a provider’s SLA must support, and whether its measurement matches what your users feel.
- When buying monitoring or observability tools. Many platforms offer SLO tracking and error-budget alerts, but definitions and data sources vary by product.
- When outsourcing operations. A managed service or SRE provider may report against SLOs; agree who defines them and how they are measured.
- When reliability and release speed collide. SLOs give a neutral basis for deciding when to slow change down.
Questions to ask vendors
- Do you publish SLOs for this service, and are they commitments with remedies or internal objectives?
- How is each indicator measured, from where, and over what window?
- Does your availability figure reflect user-facing success, or only component uptime?
- Can we see historical performance against your SLOs?
- How do your SLOs relate to the SLA credits in our contract?
- Can your platform track our own SLOs and alert on error-budget burn rate?
Our application performance monitoring and observability overview covers tools that measure SLIs and track SLOs across applications and infrastructure.
How it differs from an SLA
An SLO is a target; a service level agreement (SLA) is typically a contract. An SLA commits a provider to a level of service and usually defines what happens when it is missed, most often SLA credits, while an SLO is usually an internal goal with no financial remedy attached. Because breaking an SLA costs money, providers often set internal SLOs tighter than the SLA they sign, so they see trouble before the contract is breached. A buyer can use the same idea in reverse: set your own SLOs for what users need, then check that each provider’s SLA, measurement method and credits support them. Many SLAs measure availability alone, while good SLOs often add latency, error rates and recovery measures such as mean time to recovery (MTTR).
