Site reliability engineering (SRE) is an approach to running production systems that treats operations as a software engineering problem. SRE teams set measurable reliability targets, called service level objectives (SLOs), monitor services against them, automate repetitive operational work, and lead incident response and learning afterward. The approach was popularized by Google and is now used by many organizations that build and run their own applications. The term refers both to the practice and to the engineers, or teams, who do it.
At a glance
- SRE applies software engineering methods to keeping services reliable in production.
- Reliability is managed against measured targets (SLOs), with the gap to 100% treated as an error budget.
- SRE teams automate repetitive “toil”, run on-call and incident response, and hold blameless post-incident reviews.
- It overlaps heavily with DevOps; many describe SRE as a specific way to put DevOps ideas into practice.
- When a provider offers “SRE”, check whether it means engineering reliability improvements or mainly watching alerts.
What problem it solves
Traditionally, developers build features and an operations team keeps them running. Developers want to ship quickly; operations wants to avoid changes that break things. Without a shared measure of “reliable enough”, that tension turns into arguments, slow releases or frequent outages, and operations staff end up doing the same manual fixes again and again.
SRE addresses this by agreeing on reliability targets in advance and measuring against them. If a service is within its SLO, teams keep shipping; if the error budget is used up, they shift effort to stability. Engineers who understand both code and operations are expected to automate repetitive work rather than absorb it, so the system can grow without operations headcount growing at the same rate.
How it works
SLOs and error budgets. SRE starts by choosing indicators that reflect user experience, such as request success rate or response time, and setting targets over a period. The allowed shortfall is the error budget. Policies agree what happens when the budget runs low, such as pausing risky releases.
Monitoring and alerting. Teams instrument services with metrics, logs and traces, often through application performance monitoring and observability tools, and alert on symptoms users would notice rather than every internal fluctuation.
Reducing toil. SRE teams commonly aim to cap time spent on manual, repetitive work and use the rest for engineering: automation, better tooling and designs that remove whole classes of failure.
Incident response and learning. On-call engineers respond to incidents, aim to restore service quickly to keep mean time to recovery (MTTR) low, and run blameless post-incident reviews that focus on fixing systems and processes rather than blaming individuals.
Safe change. Gradual rollouts, automated testing and fast rollback reduce the risk that a release causes an outage, complementing formal IT change management.
Capacity and resilience. SREs plan capacity, test failover and design for high availability where the SLO requires it.
When it matters for buyers
- When you run customer-facing or revenue-critical applications. SRE practices give a structured way to manage their reliability.
- When evaluating managed cloud or application providers. “SRE” on a proposal can mean very different levels of engineering involvement.
- When outages recur. SLOs and post-incident reviews turn repeat incidents into prioritized engineering work.
- When buying observability tools. SLO tracking and error-budget alerts are common features but vary by product.
Questions to ask vendors
- What does your SRE service include: SLO design, monitoring, on-call, incident management, engineering improvements?
- Who defines the SLOs, and how are they measured and reported to us?
- What changes are your engineers allowed to make to our systems, and through what process?
- How do you run post-incident reviews, and do we receive the findings?
- How do you measure and reduce toil?
- How does your SRE team differ from a NOC or standard managed support?
Our application performance monitoring and observability overview covers tools and services that support SLO-based reliability work.
How it differs from DevOps and a NOC
DevOps is a broad movement to bring development and operations together through shared ownership, automation and fast feedback; it doesn’t prescribe a specific team structure or metrics. SRE is narrower and more prescriptive: it defines reliability through SLOs and error budgets and gives engineers explicit responsibility for production. A network operations center (NOC) is different again: it watches infrastructure, typically around the clock, and responds to alerts by following runbooks and escalating. An SRE team usually writes code to prevent problems, not just respond to them. In practice, many organizations blend all three, and the labels on provider proposals don’t always match these definitions, so judge by what the service actually does to keep uptime where it needs to be.
