Fluxtail
Log Management Guides

Service Level Objective (SLO): Definition and Examples

Learn what a service level objective means, how SLI, SLO, and SLA differ, and how to define targets, error budgets, burn-rate alerts, and examples.

By Fluxtail Engineering Updated

A service level objective (SLO) is a measurable reliability target for a service over a defined period. It says how often a user-relevant result should be good—for example, “99.9% of eligible API requests succeed over a rolling 28-day window”—and makes the measurement boundary, exclusions, and ownership explicit.

An SLO is useful only when it supports decisions. It should identify what users depend on, define exactly how good and total events are counted, create an error budget, and state what happens when reliability consumes that budget too quickly. It is not a promise that every request will work, and it is not a substitute for an SLA, monitoring, or incident response.

SLI, SLO, and SLA mean different things

The three terms describe related but separate parts of service reliability:

  • A service level indicator (SLI) is the quantitative measure. An availability SLI might be successful eligible requests divided by all eligible requests.
  • A service level objective (SLO) is the target for that indicator over a compliance period. An example is at least 99.9% success over a rolling 28 days.
  • A service level agreement (SLA) is an agreement with users or customers that includes consequences if stated commitments are not met. Credits or penalties are common examples, although consequences need not be financial.

Google's original service level terminology defines an SLO as a target value or range for a service level measured by an SLI. It distinguishes an SLA by asking what explicit consequence follows when an objective is missed. An internal objective can therefore be an SLO even when no contractual SLA exists.

Infrastructure health is not automatically an SLI. CPU utilization, queue depth, pod restarts, and disk usage can explain failures, but they usually do not measure whether a user completed an important action. Start from the user outcome, then choose the closest reliable observation point.

Start with a user journey and a service boundary

Choose one critical journey such as loading a document, submitting a payment, running a search, or receiving the output of a scheduled job. Name the service boundary responsible for the outcome. “The platform is available” is too broad; “the public checkout API accepts valid orders” can be measured.

The observation boundary matters. A backend counter cannot see a request that failed at DNS, a CDN, or a load balancer before reaching the backend. Browser instrumentation can see more of the user path, but it has its own delivery and sampling limitations. The Google SRE implementation guide separates an SLI specification—the outcome that matters—from its implementation—the telemetry and location used to measure it.

Write down both:

  • Specification: the user outcome, independent of tooling.
  • Implementation: the exact metric, log-derived count, probe, or client signal; its collection point; and known blind spots.

If the objective is about API availability at the public edge, count good and total requests at that edge. Do not take failures from the application and divide them by requests at the load balancer. Different observation boundaries can lose, add, or classify traffic differently, producing a ratio that looks precise but has no coherent meaning.

Define eligible, good, and bad events

For a request-based SLI, define three sets before choosing the target:

  1. Eligible events form the denominator. State the methods, routes, users, regions, and traffic types included.
  2. Good events form the numerator. State the exact success conditions.
  3. Bad events are eligible events that do not meet those conditions. For finalized compliance, good and bad events must partition the eligible set so that good + bad = eligible. Track temporarily unknown events separately, then define when they become good, become bad, or invalidate the provisional result.

The common form is:

SLI = good eligible events / all eligible events

Exclusions require a user-impact reason, not convenience. Decide explicitly how to treat malformed requests, client cancellations, authentication failures, load tests, health checks, internal traffic, synthetic probes, and requests made by automated retries. A 4xx response might be correct service behavior, a product failure, or an ambiguous result depending on the operation. A blanket status-code rule is rarely portable across services.

Retries need special care. Counting every attempt measures attempt reliability; counting only the final user-visible result measures outcome reliability. Neither is universally right. Record which one the SLO represents, and keep a separate metric for retry amplification because retries can conceal poor attempt reliability while adding load.

Choose request-based or windows-based measurement

A request-based SLO counts atomic events. It works well for requests, messages, records, or jobs when every event can be classified. Ten bad requests consume ten units of an event-based budget, whether they occur together or apart.

A windows-based SLO divides time into fixed windows and classifies each window as good or bad. A five-minute window might be good only when at least 99% of requests succeed and a minimum traffic condition is met. Google Cloud's service monitoring concepts distinguishes request-based SLIs from windows-based SLIs that count periods meeting a goodness criterion.

The two models answer different questions. Event-based measurement weights busy periods more heavily because more users are affected. Window-based measurement weights each time window equally and can represent “good user minutes” or scheduled checks. State the window size and classification rule; changing either changes the meaning of the SLO.

Choose a rolling or calendar compliance period

A rolling window always examines the most recent duration. It does not forget an outage at a month boundary, and it is useful for ongoing operational decisions. A calendar window resets on a business boundary such as a calendar month or quarter, which can align with reporting and planning.

Google's guidance on choosing an SLO window notes that rolling windows follow user experience more closely, while calendar windows align more naturally with business planning. For rolling windows, an integral number of weeks avoids changing the weekday-to-weekend mix. Record the period type, duration, time zone for calendar boundaries, and evaluation cadence.

Calculate the SLO and error budget correctly

For an event-based objective:

compliance = good eligible events / all eligible events
allowed bad-event fraction = 1 - SLO target
error budget = eligible events × (1 - SLO target)

Suppose an API receives 10,000,000 eligible requests during its compliance period and has a 99.9% availability SLO:

target              = 0.999
allowed bad fraction = 1 - 0.999 = 0.001
error budget         = 10,000,000 × 0.001 = 10,000 bad requests

If 2,500 eligible requests are bad, the service has consumed 25% of its request error budget. It still has 7,500 bad requests of budget remaining for that period.

Do not convert those 10,000 bad requests into “minutes of downtime.” Request volume changes over time, and an event-based SLO weights failures by affected requests. A downtime allowance is valid only for an explicitly time- or window-based SLI. Even then, define whether a minute is bad after one failed probe, a threshold of failed requests, or another condition.

The target is a product and engineering choice, not a universal maturity score. A stricter target costs more error budget for the same failure rate and may demand more engineering effort. Choose it from user expectations, business harm, dependency limits, current evidence, and the actions the organization is prepared to take. Google's SRE guidance recommends starting with a small number of representative objectives and refining them rather than treating 100% as the default.

Measure latency as threshold compliance

“Average latency below 300 ms” can hide a slow minority of requests. An SLO is clearer when it counts the proportion of eligible requests below a user-relevant threshold:

good latency events = eligible requests completed in 300 ms or less
latency SLI = good latency events / all eligible requests

An objective might require 95% of eligible requests to finish within 300 ms over a rolling 28 days. This is not the same as saying p95 is always below 300 ms in every short interval. Define the aggregation and compliance period exactly. If two thresholds protect different experiences, write two objectives, such as a common-response threshold and a slower tail threshold, rather than averaging them together.

Use a complete service level objective template

The following template is deliberately plain. A monitoring rule or SLO tool can implement it later, but stakeholders should be able to review the definition without reading a query language.

Name and version:
Owner and approvers:
User journey:
Service boundary:

SLI specification:
SLI implementation and observation point:
Eligible events:
Good-event rule:
Bad- and unknown-event handling:
Explicit exclusions and reasons:

Target:
Compliance period: rolling or calendar, duration, time zone if relevant
Evaluation cadence:
Data delay and completeness limit:

Error-budget calculation:
Burn-rate alerts:
Low-traffic safeguard:
Error-budget policy and owner:

Dashboard and raw evidence links:
Approval date:
Next review date:
Change history:

Version the SLO definition alongside its query or configuration. A schema, route, status mapping, or traffic exclusion can change compliance without changing user experience. A versioned definition makes that discontinuity visible instead of silently rewriting history.

Four practical SLO examples

These examples show the shape of an objective, not universal targets. Replace every threshold and classification with values justified for the service.

API availability

Objective: At least 99.9% of eligible requests to the public order-submission endpoint succeed over a rolling 28-day period, measured at the public load balancer.

Eligible: Authenticated requests with a syntactically valid order, excluding documented synthetic checks and approved load tests.

Good: The edge records the application-specific accepted or idempotently-replayed result. Define timeout and ambiguous upstream outcomes as bad until reconciliation proves otherwise.

This example avoids assuming that every 4xx is good or every retry is a new user action. An idempotency key can help relate repeated attempts, but the SLI implementation must say whether it counts attempts or final operations.

API latency

Objective: At least 95% of eligible catalog reads complete at the public edge within 300 ms over a rolling 28-day period.

Good: An eligible request has a completed edge duration at or below the threshold. Decide whether failed requests are also latency-eligible; excluding failures from a latency denominator can make a broken service appear fast.

Segment reporting by route class or region when those experiences differ, but avoid dimensions so specific that each series has too little traffic to evaluate. Keep the contractual SLO definition stable while using diagnostic dimensions to explain misses.

Async job freshness and completion

Objective: At least 99% of eligible hourly settlement runs publish a validated result no later than 20 minutes after their scheduled time during a calendar month.

Eligible: Scheduled production runs after deduplicating scheduler retries by a stable run identifier.

Good: Exactly one accepted result reaches the published state before the deadline and passes the defined validation. A started job is not a completed job. A duplicate success may still be bad if it creates duplicate user-visible output.

This is event-based even though it refers to time. A windows-based alternative could classify each scheduled hour as good or bad. Pick one model and do not combine a job count from the scheduler with completions counted at a different pipeline stage unless the identifiers reconcile completely.

Ingestion completeness and correctness

Objective: At least 99.5% of eligible source records are present once, with required fields valid, in the destination by the agreed deadline over a rolling four-week period.

Eligible: Source records committed within the period, identified at the authoritative source boundary.

Good: A source record has one accepted destination record, required identifiers match, and the record arrived by the deadline. Completeness and correctness may deserve separate SLIs if one query cannot express both without hiding the failure mode.

The denominator cannot come only from the destination: missing records are invisible there. Use a source count or reconciled manifest, account for late data, and document how corrections that arrive after the deadline affect the original period.

Protect the measurement from telemetry errors

An SLO can fail as a measurement before the service fails. Treat the telemetry path as a dependency and test these conditions:

  • Missing events: Exporter or ingestion loss can remove good and bad events. Monitor expected volume and collection freshness separately.
  • Duplicates: At-least-once delivery can count one outcome more than once. Deduplicate with a stable event or operation identifier where the SLI requires unique outcomes.
  • Sampling: Sampled events cannot be treated as an exact denominator without a statistically valid design. Never sample away failures selectively and then call the result availability.
  • Delayed data: A late event may revise compliance after an alert or report was generated. Publish the allowed data-latency window and whether results are provisional.
  • Schema changes: Renaming a route, status, or success field can produce a false improvement or collapse. Contract-test the classification query before rollout.
  • Clock errors: Event time, observation time, and scheduled time have different meanings. Preserve the source timestamp and use synchronized clocks where deadlines depend on elapsed time.
  • Unknown results: Do not silently discard records that cannot be classified. Track them and decide whether they are bad, temporarily unknown, or grounds for invalidating the measurement.

Compare SLO drops with known incidents and user reports. An objective that remains green during a known user-impacting failure is measuring the wrong boundary or classification. An objective that turns red without observable harm may still reveal a real risk, but it deserves validation before the target drives policy.

Alert on error-budget burn, not every brief error

An error budget turns a target into an operating rate. For a 99.9% objective, the sustainable bad-event rate is 0.1%. Burn rate is the observed bad-event rate divided by that sustainable rate:

burn rate = observed bad-event rate / (1 - SLO target)

A burn rate of 1 means the current failure rate would exactly consume the budget if sustained for the whole compliance period. A rate above 1 would exhaust it; a rate below 1 consumes it more slowly. Google Cloud documents the same normalized interpretation in its burn-rate alerting guide.

Use multiwindow, multi-burn-rate alerts to distinguish fast incidents from slow deterioration. A long window shows that enough budget has been consumed to matter; a paired short window confirms the burn is still active. Separate high-burn paging from lower-burn ticket or review paths.

Google's SLO alerting chapter gives, for a 99.9% objective over a 30-day compliance period, starting examples such as a 14.4× burn over both one hour and five minutes, and a 6× burn over both six hours and 30 minutes. Those values are examples, not constants to paste into every service. Adapt them to the target, compliance period, traffic, response time, paging capacity, and impact of a failed event. Test detection and reset behavior with recorded or synthetic failures before relying on the alerts.

Add safeguards for low traffic

Ratio alerts become unstable when the denominator is small. One failed request out of ten produces a 10% error ratio even if there is not enough evidence to distinguish a persistent outage from one isolated event. Conversely, waiting for a statistically larger sample may delay response to a rare but valuable operation.

Choose a low-traffic rule from user impact:

  • combine the ratio with a minimum eligible-event or absolute bad-event count;
  • use a synthetic check for a critical path, but keep synthetic traffic separate from the user SLI unless the definition includes it;
  • group closely related services only when they represent one journey and do not hide a complete failure of a small component;
  • alert directly on a single failure when one failed event has severe business impact;
  • use a longer-window review or ticket when immediate paging would create noise without changing the response.

Google's low-traffic guidance warns that ordinary burn-rate alerts need adjustment when few requests arrive. The objective is not to make the graph quiet; it is to preserve meaningful detection for the service's actual impact.

Give the error budget an explicit policy

An error budget without a policy is only a report. Write the actions, decision owners, exceptions, and exit conditions before the service exhausts its budget.

A practical policy should state:

  • who owns the SLO and who can approve changes to it;
  • which burn alerts page, create a ticket, or start a review;
  • what happens as remaining budget falls through agreed thresholds;
  • whether risky releases pause after budget exhaustion and which security or emergency changes remain allowed;
  • how third-party failures, planned tests, and misclassified events are reviewed;
  • what evidence permits normal release activity to resume;
  • how disagreements about classification or policy are escalated.

The published Google example error-budget policy is useful as a structure, but its thresholds and actions belong to its example service. Copying the numbers without obtaining agreement from product, engineering, and operations leaves the policy without authority. The policy should enable reliability work, not punish an individual or team for reporting failures.

Review the objective on a declared schedule and after material changes to the journey, architecture, traffic, telemetry, or business impact. Record the author, reviewers, approvers, rationale, query version, and next review date. Keep prior definitions so reports can explain why two periods are not directly comparable.

Use metrics for SLO computation and logs for evidence

A metrics or dedicated SLO system is normally the right place to compute ratios, rolling windows, budgets, and burn-rate alerts. It can maintain consistent counters and evaluate alert windows efficiently. Logs can support that system when every eligible event is delivered, classified, and deduplicated under a controlled schema, but that assumption must be tested.

Logs are especially useful after an SLO signal changes. They can show which route, release, region, dependency result, or error class contributed to bad events. Preserve stable identifiers and fields such as the SLO classification version, service, environment, route template, outcome, duration bucket, release, and request or operation ID. Do not put tokens, credentials, raw request bodies, or unnecessary personal data into the event.

For broader collection and protection practices, see log management best practices. For the separate question of how long incident handling takes, see mean time to resolution; MTTR does not replace an SLO and an SLO does not measure every incident phase.

Fluxtail is a paid, self-service Starter/Pro logs service, not an SLO calculator, metrics dashboard, tracing/APM system, or incident-management platform. After a supported sender or collector delivers events and the field mapping is verified, its Live Tail and search filters can help inspect retained rows behind an SLO change. The stored field names depend on the source and receiver mapping, so verify one raw row before relying on a filter.

Fluxtail's built-in AI chat and hosted MCP are separate interfaces. The hosted MCP uses account-bound OAuth with PKCE; read operations remain limited to the authorized account, and mutations are proposed before a short-lived confirmation is applied. Use either interface to narrow an investigation, then confirm conclusions against the retained raw events. Neither interface turns log evidence into an authoritative SLO denominator.

If centralized log evidence is the missing part of the investigation path, create a Fluxtail account and choose the paid plan that matches the required ingest and retention.

Validate the SLO before it governs decisions

Before enforcing a new objective, run it in observation mode and verify:

  1. The user journey and service boundary are named clearly.
  2. Good, bad, eligible, excluded, and unknown events are exhaustive and reviewable.
  3. Numerator and denominator come from the same observation point and population.
  4. Retry, cancellation, synthetic, internal, and client-error behavior is explicit.
  5. The target and period reflect user and business needs rather than a copied industry number.
  6. The error-budget arithmetic matches the SLI type.
  7. Missing, duplicate, sampled, late, and schema-invalid telemetry is visible.
  8. Known incidents consume budget in a way that resembles user impact.
  9. Burn-rate alerts detect seeded fast and slow failures and reset after recovery.
  10. Low-traffic behavior has an intentional rule.
  11. The error-budget policy has owners, actions, exceptions, and exit conditions.
  12. The definition, query, approvals, and review date are versioned together.

A good service level objective is compact enough to explain, precise enough to implement twice with the same result, and important enough to change a real decision. Start with one critical journey, validate its measurement, and add another objective only when it protects a distinct user outcome.