Observability

Health that explains why · SLOs with error budgets · incidents that arrive pre-diagnosed with their support bundle attached · alerts routed to owners.
Review plan-9107 — waiting 11h

SLO — Platform availability

30d rolling
99.94% target 99.9%
Error budget remaining
62% left · burn rate 0.4× · healthy

SLO — API latency p99

30d rolling
212 ms target < 400 ms

SLO — Upgrade success

rolling 90d · fleet
96% 24 of 25 verified
1 rolled back automatically (helm timeout) — halted cleanly, bundle attached, zero data loss.

Fleet health matrix

why, on hover — not just red/green
EnvironmentAPIWorkloadsDataIngressBackups
meridian-health/prod okrollingokokok
kestrel-bank/prod okdegradedokcert 12dok
northwind-logistics/prod okokokokok
verde-foods/prod okokokokok
cobalt-mining/prod air-gapped — snapshot from signed bundle, 3d old · all green at export

Active alerts

2 firing · 14 environments monitored
AlertEnvironmentSinceSeverity
SearchReadinessLow
linked to INC-241 — see the Incidents tab
kestrel-bank/prod03:12critical
CertExpirySoon
api.kestrel.example · 12d
kestrel-bank/prod2dwarning
Alert rules ship with the release and are tuned as config plans — a noisy rule is a diff, not a dashboard tweak.

Capacity & forecast

fleet aggregate · 90d
vCPU utilisation
68%
Memory utilisation
71%
Storage growth
41.2 TB
GPU allocation
14 / 18
Forecast: kestrel-bank memory ceiling in ~6 weeks; quill-legal/dev idle 21 days (downscale suggested, ~$610/mo). Recommendations become planned operations — reviewed and approved, never auto-applied.
Open
1 sev-2
INC-241 · kestrel-bank / prod
MTTA · 30d
2m 10s
alert → acknowledged
MTTR · 30d
38m
−22m vs June — diagnosis arrives with the page
Auto-resolved before paging
71%
self-heal loop: detect → diagnose → fix → retry

INC-241 — Search degraded

investigating kestrel-bank / prod · sev-2 · owner: Priya N.
03:12
Detected: platform-search readiness < 2/3 for 5m · alert SearchReadinessLow
03:12
Bundle captured at detection — logs, k8s events, values (secrets masked), journal · bundle-inc-241.tgz
03:13
Agent diagnosis: OOMKilled — ingest batches 2× larger since 0.1.90, memory limit unchanged. Evidence cited: k8s-events#88, journal#412
03:13
Self-heal attempted: restart with backoff — recurred twice → escalated with one precise question: “Raise search memory 2Gi → 4Gi? (+$96/mo)”
03:14
Paged Priya N. · acknowledged in 1m 40s · war-room opened
03:20
Known-issue match: failure signature seen at 2 other sites since 0.1.90 — permanent fix (dynamic batch cap) already staged for 0.1.92
08:40
Awaiting approval — remediation plan plan-9107: memory 2Gi → 4Gi now, batch cap at 0.1.92. Nothing hot-patches production.
Review plan-9107 in Approvals Case CASE-1042 Copilot analysis

Recent incidents

IDSummaryEnvironmentSevDurationStatus
INC-241Search degraded — OOMKilledkestrel-bank/prod211h openinvestigating
INC-238Cert renewal stuck (DNS-01 throttle) — self-healed on retryhelios-energy/prod318mresolved
INC-233vCPU quota exhausted mid-scale — preflight gap, now checkedtern-aviation/staging342mresolved
INC-229Upgrade rolled back cleanly (helm timeout) — zero data lossnorthwind-logistics/prod231mpostmortem
INC-224Ingress cert expiry near-miss — renew automation fixedverde-foods/prod425mpostmortem

Postmortems & learning

blameless · actions tracked as planned operations
Published · 90d
4 · blameless template · action items tracked as planned operations
Actions closed
9 of 11 — 2 in flight (batch cap → 0.1.92, quota preflight shipped in 0.1.91)
Fed back to product
each incident emits a structured failure signature — 3 became engine fixes, so the next customer never sees them
Incident → signature → engine fix → fleet-wide: INC-233's quota gap became a preflight check, shipped in 0.1.91.

Incidents by month

fleet · paged incidents only

On-call now

follow-the-sun · platform SRE
PN
Priya Natarajan
primary · owns INC-241 · until 18:00 CET
engaged
KT
Kai Tran
secondary · escalation after 15 min unacked

Notification routing

SeverityChannelEscalation
criticalpage on-call + #ops-alerts15 min → secondary · 30 min → manager
warning#ops-alerts (Slack) + emailbusiness hours triage
infoweekly digest
operation eventsinitiator + tenant adminsfailed op ⇒ warning route
Channels: PagerDuty, Slack, email, webhooks. Air-gapped sites queue notifications into the signed bundle export. Sev-1/2 auto-open a vendor case with the bundle attached (per your approval policy).

Paging health · 30d

Pages
6 −40% MoM
Median ack
2m 10s
Auto-diagnosed
6 of 6 arrived with bundle + suggested runbook
False pages
1 · alert rule tuned (PR linked)
Postmortems
2 published · actions tracked as planned operations
pages/month — every support call that doesn't happen