1.2 KiB
1.2 KiB
Sprint 68 Plan: Corpus Expansion and Gold Standard Benchmarks
Context
Long-term quality depends on a representative corpus. Sprint 68 creates a broad, versioned benchmark corpus across paradigms and pair classes.
Goals
- Build and version gold corpora by language family
- Add benchmark tiers (unit, module, service, legacy project)
- Support reproducible benchmark replay
- Publish benchmark-driven pair scorecards
Steps
Step 909: Corpus manifest schema v1 (12 tests)
Step 910: Corpus ingestion tooling and normalization (10 tests)
Step 911: Gold benchmark curation pipeline (10 tests)
Step 912: Replay runner with deterministic seeds/artifacts (10 tests)
Step 913: Benchmark scorecard generator per pair (10 tests)
Step 914: Corpus version diff and migration tools (8 tests)
Step 915: whetstone_run_pair_benchmark MCP tool (8 tests)
Step 916: whetstone_get_pair_scorecard MCP tool (8 tests)
Step 917: Benchmark publication bundle (8 tests)
Step 918: Sprint 68 integration summary + regression (8 tests)
Benchmark Rule
- Stable-tier promotions require passing the current gold benchmark set.