WIP: stage all uncommitted work — sprints 46-221, graduation headers, specialist fleet, test steps 909-1988
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
36
sprint68_plan.md
Normal file
36
sprint68_plan.md
Normal file
@@ -0,0 +1,36 @@
|
||||
# Sprint 68 Plan: Corpus Expansion and Gold Standard Benchmarks
|
||||
|
||||
## Context
|
||||
|
||||
Long-term quality depends on a representative corpus. Sprint 68 creates a broad,
|
||||
versioned benchmark corpus across paradigms and pair classes.
|
||||
|
||||
---
|
||||
|
||||
## Goals
|
||||
|
||||
1. Build and version gold corpora by language family
|
||||
2. Add benchmark tiers (unit, module, service, legacy project)
|
||||
3. Support reproducible benchmark replay
|
||||
4. Publish benchmark-driven pair scorecards
|
||||
|
||||
---
|
||||
|
||||
## Steps
|
||||
|
||||
### Step 909: Corpus manifest schema v1 (12 tests)
|
||||
### Step 910: Corpus ingestion tooling and normalization (10 tests)
|
||||
### Step 911: Gold benchmark curation pipeline (10 tests)
|
||||
### Step 912: Replay runner with deterministic seeds/artifacts (10 tests)
|
||||
### Step 913: Benchmark scorecard generator per pair (10 tests)
|
||||
### Step 914: Corpus version diff and migration tools (8 tests)
|
||||
### Step 915: `whetstone_run_pair_benchmark` MCP tool (8 tests)
|
||||
### Step 916: `whetstone_get_pair_scorecard` MCP tool (8 tests)
|
||||
### Step 917: Benchmark publication bundle (8 tests)
|
||||
### Step 918: Sprint 68 integration summary + regression (8 tests)
|
||||
|
||||
---
|
||||
|
||||
## Benchmark Rule
|
||||
|
||||
- Stable-tier promotions require passing the current gold benchmark set.
|
||||
Reference in New Issue
Block a user