37 lines
1.3 KiB
Markdown
37 lines
1.3 KiB
Markdown
|
|
# Sprint 77 Plan: SLA Reliability and Incident-Ready Porting Operations
|
||
|
|
|
||
|
|
## Context
|
||
|
|
|
||
|
|
At scale, transpilation is an operational system with uptime/reliability targets.
|
||
|
|
Sprint 77 defines SLAs and incident workflows for porting infrastructure.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Goals
|
||
|
|
|
||
|
|
1. Define and enforce transpilation service SLAs/SLOs
|
||
|
|
2. Add reliability signals and alerting for pair regressions
|
||
|
|
3. Integrate incident runbooks for gate and certification failures
|
||
|
|
4. Improve recovery time with automated containment strategies
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Steps
|
||
|
|
|
||
|
|
### Step 999: Transpilation SLA/SLO schema (`availability`, `latency`, `quality`) (12 tests)
|
||
|
|
### Step 1000: Reliability signal collector integration (10 tests)
|
||
|
|
### Step 1001: Alert policy engine for SLO breaches (10 tests)
|
||
|
|
### Step 1002: Incident classifier for transpilation failures (10 tests)
|
||
|
|
### Step 1003: Auto-containment strategy hooks (10 tests)
|
||
|
|
### Step 1004: Incident runbook bindings and drill tracker integration (8 tests)
|
||
|
|
### Step 1005: `whetstone_get_transpilation_slo_status` MCP tool (8 tests)
|
||
|
|
### Step 1006: `whetstone_trigger_porting_incident_drill` MCP tool (8 tests)
|
||
|
|
### Step 1007: Reliability operations report artifact (8 tests)
|
||
|
|
### Step 1008: Sprint 77 integration summary + regression (8 tests)
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Reliability Rule
|
||
|
|
|
||
|
|
- Repeated SLO breach for a pair automatically downgrades support tier pending review.
|