Files
whetstone_DSL/sprint68_plan.md

1.2 KiB

Sprint 68 Plan: Corpus Expansion and Gold Standard Benchmarks

Context

Long-term quality depends on a representative corpus. Sprint 68 creates a broad, versioned benchmark corpus across paradigms and pair classes.


Goals

  1. Build and version gold corpora by language family
  2. Add benchmark tiers (unit, module, service, legacy project)
  3. Support reproducible benchmark replay
  4. Publish benchmark-driven pair scorecards

Steps

Step 909: Corpus manifest schema v1 (12 tests)

Step 910: Corpus ingestion tooling and normalization (10 tests)

Step 911: Gold benchmark curation pipeline (10 tests)

Step 912: Replay runner with deterministic seeds/artifacts (10 tests)

Step 913: Benchmark scorecard generator per pair (10 tests)

Step 914: Corpus version diff and migration tools (8 tests)

Step 915: whetstone_run_pair_benchmark MCP tool (8 tests)

Step 916: whetstone_get_pair_scorecard MCP tool (8 tests)

Step 917: Benchmark publication bundle (8 tests)

Step 918: Sprint 68 integration summary + regression (8 tests)


Benchmark Rule

  • Stable-tier promotions require passing the current gold benchmark set.