Upgrade 0.1.88 → 0.1.91 Running

meridian-health / prod · azure westeurope · operation run-8842 · plan pl-9f2e41 · started 14:02 by kai.tran · journaled & resumable

Workflow run view

click a node to inspect
verified running blocked (checkpoint) pending / skipped by when: failed
Point of no return crossed at 14:19. Everything before migrations was fully rollback-able. From here, database migrations and infra mutations are forward-only — on failure the engine halts with a support bundle and offers resume or restore.

Upgrade verified — 0.1.88 → 0.1.91

verified 14:44 · 42 minutes end-to-end
Verify
green deployment readiness · endpoint health · cert validity · version fingerprints
Smoke tests
6 / 6 product postUpgrade hooks
Duration
42m — preflight 14:02 → verified 14:44 (incl. 11m waiting on the point-of-no-return approval)
Vendor involvement
0 hours — run end-to-end by kai.tran
Rollback window
snapshot pl-9f2e41-snap retained 14d — Helm/config rollback stays one click
Back to Operations

Rollback failed — platform-registry

halted · state fully journaled · nothing re-tried blindly
Revert of platform-registry (rev 21 → 20) timed out — a PVC resize was still in flight. The engine stopped at the first rollback failure; it did not attempt the remaining reverts.
Reverted
3 releases back on recorded 0.1.88 revisions
Not attempted
8 releases still on 0.1.91 — healthy, untouched
Failed revert
platform-registry rev 21 → 20 · timeout 300s
Data
migrations are expand-first — the schema serves both versions; no data was touched by rollback
Environment now
serving traffic mixed versions · search degraded · nothing half-applied inside any single release
Escalation
vendor case CASE-1043 auto-opened (sev-2, per policy) · bundle bundle-run-8842b.tgz attached · on-call paged
Case CASE-1043
The journal records exactly what reverted and what didn't — per release.

Operation failed — helm_rollout

halted cleanly · nothing half-applied
Helm upgrade timeout: release platform-search stuck — 2 of 3 pods failed readiness within 300s (OOMKilled). The original error is preserved.
Support bundle auto-attached: bundle-run-8842.tgz (34 MB) — logs, events, Kubernetes resources, Helm values, versions, operation journal, secrets masked.
Rollback available for the rollback-able set: Helm releases → recorded prior revisions, configuration → prior versioned copy, manifest pointer. Migrations already ran expand-first and are backward-compatible, so the 0.1.88 pods run safely against the migrated schema.

Live log

streaming over the event channel · journal spans: workflow → task → operator → command

Inspector — helm_rollout

Release gate

checked at preflight
Target release
0.1.91 signed
Signature
cosign · key converge-release-2026 · verified offline
Upgradable from
≥ 0.1.1 · this env: 0.1.88 ✓
Kubernetes
requires ≥1.30 · cluster: 1.31.7 ✓
Cloud / profile
azure ✓ · production ✓
Backup taken
3 databases + tfstate + state + secret inventory · 14:07

Rollback semantics

rollback-able Helm releases · configuration/overlays · manifest pointer
forward-only database migrations (expand-first) · infrastructure mutations

Preflight must pass and a snapshot must exist before the point of no return may be crossed. If rollback itself fails, the operation lands in RollbackFailed with a bundle attached.

Demo