# Sprint 68 Plan: Corpus Expansion and Gold Standard Benchmarks ## Context Long-term quality depends on a representative corpus. Sprint 68 creates a broad, versioned benchmark corpus across paradigms and pair classes. --- ## Goals 1. Build and version gold corpora by language family 2. Add benchmark tiers (unit, module, service, legacy project) 3. Support reproducible benchmark replay 4. Publish benchmark-driven pair scorecards --- ## Steps ### Step 909: Corpus manifest schema v1 (12 tests) ### Step 910: Corpus ingestion tooling and normalization (10 tests) ### Step 911: Gold benchmark curation pipeline (10 tests) ### Step 912: Replay runner with deterministic seeds/artifacts (10 tests) ### Step 913: Benchmark scorecard generator per pair (10 tests) ### Step 914: Corpus version diff and migration tools (8 tests) ### Step 915: `whetstone_run_pair_benchmark` MCP tool (8 tests) ### Step 916: `whetstone_get_pair_scorecard` MCP tool (8 tests) ### Step 917: Benchmark publication bundle (8 tests) ### Step 918: Sprint 68 integration summary + regression (8 tests) --- ## Benchmark Rule - Stable-tier promotions require passing the current gold benchmark set.