← Back to catalogue
TODO

Service Level Objective / Reliability Commitment

vr.wm-sft-016 · wm-sft-016-service-level-objective-reliability-commitment

Immutable SLI and SLO policies evaluations error-budget rules and observability bindings

World Models Information and virtual systems INF.SFT.SLO

Bundle → Layer → Finding → Questions Filled

3 bundles · 3 layers · 5 findings · 11 questions

Indicator definition What is measured and how.

Service level indicator

The measurement the objective is based on.

Indicator specification

The events counted as good and valid, and the data source.

  1. Which indicator does this objective use, and how are good and valid events defined?
  2. From which telemetry source and query is the indicator computed?

Measurement scope

The service, journey and user population covered.

  1. Which service and user journey does the indicator cover, and for which users?
  2. Which traffic is excluded from the measurement, and on what grounds?
Objective and budget The target and what happens when it is missed.

Target

Threshold, window and version of the objective.

Threshold and window

The target value and the period it is evaluated over.

  1. Which target value and evaluation window does this objective set?
  2. Which version of the objective is in force, and since when?

Error budget rules

The budget derived from the target and the actions it triggers.

  1. How much error budget does the target allow in the window, and how much is consumed?
  2. Which actions follow when the budget is exhausted or burning fast?
Evaluation and accountability How results are recorded and who answers for them.

Evaluation record

Results per window and the owning team.

Evaluation result

Whether the objective was met in a window.

  1. Was the objective met in the last completed window, and by what margin?
  2. Which team owns the objective, and which service does it protect?
  3. Which incidents consumed the budget in this window?

Classifiers Filled

Family
World Models
Category
Information and virtual systems
Entry kind
standalone-mm
Navigation path
NAV.INF.SFT.SLO
Domain
INF.SFT.SLO
Industry
Cross-industry
Tags
servicelevelobjectivereliabilitycommitmentinf.sft.slo

What it is Filled

A service level objective is a stated reliability target for a software service, defined as a service level indicator measured over a window with a threshold, together with the error budget rules that follow from it. The objective is the policy record; the telemetry that measures it and the contractual agreement with customers are separate things.

Why it exists Filled

Immutable SLI and SLO policies evaluations error-budget rules and observability bindings

Distinguishing features Filled

  • The objective is an internal reliability policy; a service level agreement is a contract with consequences for customers.
  • The indicator definition is part of the objective, so changing the query changes the objective.
  • Error budget rules link measured reliability to decisions about shipping and remediation.
  • Distinct from the telemetry signal it is computed from and from the incident records that consume the budget.

What robots and AI may and may not do Filled

Must not

  • Change a threshold, window or indicator query to make a missed objective look met.
  • Exclude traffic from the measurement without recording the exclusion and its reason.
  • Present an internal objective as a contractual commitment to customers.
  • Report results from an incomplete window as if the window had closed.
  • Suppress budget burn alerts to protect a release schedule.

Only with a human decision

  • Approving a new objective or a change to a threshold, window or indicator.
  • Deciding to halt or continue releases when the error budget is exhausted.
  • Communicating reliability results to customers or regulators.

May

  • Compute indicator values and error budget consumption from the declared telemetry source.
  • Report objective results per window with the data source and query version.
  • Raise alerts when budget burn exceeds the declared rules.
  • Propose objective changes with the evidence behind them.

Moral aspects Filled

  • Reliability targets decide which users' failures are treated as acceptable, and averages can hide groups that always suffer.
  • Error budget pressure falls on on-call staff, whose working conditions are part of the cost.
  • Services people depend on for health, money or safety deserve targets set with that dependence in mind.

Who is affected

  • Users of the service
  • On-call engineers and support staff
  • Customers relying on contractual commitments
  • Teams owning dependent services

Owners Filled

Steward

The team that operates the service and answers for its reliability targets.

Master systems

  • Observability and monitoring platform
  • Service catalogue
  • Incident management system

Links to other meta-models Filled

references

  • wm-sft-017-telemetry-operational-signal - The indicator is computed from telemetry signals.

part-of

  • wm-sft-001-software-product - An objective protects a service of a software product.

neighbor

  • world-k4-service - Contractual service level commitments to requesters live in the service model.

What else AI and robots need to interact with it Filled

Identity and identifiers required Filled

  • An objective is identified by a name scoped to its service in the service catalogue, with a version.
  • The indicator definition cites the telemetry query and its version as part of that identity.

Direct properties not applicable Not applicable

Not applicable

An objective is a policy record, not a physical object; its threshold, window and budget figures are part of its structure.

Recognition optional Filled

  • Marked by an indicator definition, a numeric target, an evaluation window and an owning team.
  • Often confused with a service level agreement, with an alert threshold or with a raw metric.

Capabilities and actions required Filled

  • Budget consumption and burn rate can be computed continuously within a window.
  • Results can be compared across windows and services.
  • Objective versions can be compared to show what changed and when.

Hazards and failure modes required Filled

  • Targets quietly weakened to avoid reporting a miss.
  • Indicators that measure the system rather than the user experience.
  • Alert fatigue from burn rules set without regard to on-call load.
  • Objectives treated as contractual promises they were never meant to be.

Standards and interfaces required Filled

  • OpenTelemetry for metrics, traces and logs feeding the indicator.
  • Prometheus and OpenMetrics exposition for metric collection.
  • ISO/IEC 20000-1 where objectives feed a service management system.

Context of use required Filled

  • Used by platform and product engineering teams to steer reliability work and release decisions.
  • Objectives for regulated or critical services may have to be shown to supervisors.

Sources Filled

  1. OpenTelemetry Specification, Cloud Native Computing Foundation
  2. ISO/IEC 20000-1 Service management system requirements, ISO and IEC
  3. OpenMetrics specification, Cloud Native Computing Foundation

Open questions

  • Planned model: boundary questions, research and every section remain to be written.

Machine files

Provenance

planned (registry candidate) · todo

Built from: models/runtime-index.json, ver-cy/world-models/card-supplements/wm-sft-016-service-level-objective-reliability-commitment.json

Planned entry, hidden from the catalogue until researched.