WIP: stage all uncommitted work — sprints 46-221, graduation headers, specialist fleet, test steps 909-1988

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Bill
2026-04-22 10:15:48 -06:00
parent 486940cbe4
commit 72ffee68fa
2179 changed files with 82979 additions and 1150 deletions

36
sprint68_plan.md Normal file
View File

@@ -0,0 +1,36 @@
# Sprint 68 Plan: Corpus Expansion and Gold Standard Benchmarks
## Context
Long-term quality depends on a representative corpus. Sprint 68 creates a broad,
versioned benchmark corpus across paradigms and pair classes.
---
## Goals
1. Build and version gold corpora by language family
2. Add benchmark tiers (unit, module, service, legacy project)
3. Support reproducible benchmark replay
4. Publish benchmark-driven pair scorecards
---
## Steps
### Step 909: Corpus manifest schema v1 (12 tests)
### Step 910: Corpus ingestion tooling and normalization (10 tests)
### Step 911: Gold benchmark curation pipeline (10 tests)
### Step 912: Replay runner with deterministic seeds/artifacts (10 tests)
### Step 913: Benchmark scorecard generator per pair (10 tests)
### Step 914: Corpus version diff and migration tools (8 tests)
### Step 915: `whetstone_run_pair_benchmark` MCP tool (8 tests)
### Step 916: `whetstone_get_pair_scorecard` MCP tool (8 tests)
### Step 917: Benchmark publication bundle (8 tests)
### Step 918: Sprint 68 integration summary + regression (8 tests)
---
## Benchmark Rule
- Stable-tier promotions require passing the current gold benchmark set.