Skip to content
iam.alan.abreu
← Work
ONGOINGDraft2023-09present

Reliability & Observability Baseline

A shared SLO, alerting and instrumentation baseline that made service health legible across teams.

Domains
  • Reliability
  • Cloud
  • Platform
Technologies
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Loki
  • Alertmanager

Overview

One instrumentation and alerting contract every service inherits, so on-call does not require service-specific tribal knowledge.

Context

Each service had its own dashboards, its own alert thresholds and its own definition of healthy.

Problem

Incidents were slow to diagnose because responders had to learn the service before they could read its signals.

Constraints

Cardinality and retention cost had to stay bounded, and instrumentation could not require rewriting existing services.

Architecture

OpenTelemetry instrumentation feeds a shared metrics and logs stack. SLOs are declared next to the service and rendered into dashboards and alerts automatically.

Decisions

Alerted on symptoms defined by SLO burn rate rather than on resource-level thresholds.

Implementation

Started with the three services that generated the most pages, then generalised the pattern.

Results

Placeholder — measured outcomes will be published with the real case study.

Lessons

Deleting alerts improved reliability more than adding them.

Linked from

Related suggestions