Files
whetstone_DSL/docs/progress_log_2026-02-26.md

19 KiB

Progress Log - 2026-02-26

Scope Covered Today

  • Sprint tranche execution and validation from 206 through 227.
  • Hard benchmark reruns on:
    • challenging_subset_prod_2026-02-26.jsonl
    • challenging_fullstack_multifile_2026-02-26.jsonl
  • Production-loop anti-false-green hardening, projection/fullstack contract enforcement, parity gating, and parity remediation.

Key Outcomes

  1. Sprints 206-211:
  • Added anti-false-green production evidence gating.
  • Added projection-contract validation and parity/efficiency/intent consistency analyzers.
  • Status: PARTIAL (diagnostics/enforcement improved, parity and readiness still weak at that stage).
  1. Sprints 217-221:
  • Implemented multi-file/fullstack contract gates (migration/security/SLO/rollout checks).
  • Added fullstack closure analyzer and execution tracker.
  • Status: IMPLEMENTED, then PARTIAL until parity inconsistency was addressed.
  1. Sprints 222-224:
  • Added strict AB/production parity hard gate and summary metrics.
  • Eliminated unresolved divergence by blocking inconsistent runs.
  • Status: DONE (for gating/reporting behavior).
  1. Sprints 225-227:
  • Added deterministic path-B repair layer:
    • C++ include repair for std::vector.
    • Go/Rust queue transpile skeleton normalization.
  • Integrated repair into AB path-B flow.
  • Fixed Rust gate harness initialization issue.
  • Validation result on latest reruns:
    • challenging_subset_prod_20260226_r7: ab_prod_divergence_count=0, ab_consistency_blocked_count=0
    • challenging_fullstack_multifile_20260226_r7: ab_prod_divergence_count=0, ab_consistency_blocked_count=0
  • Status: DONE.

Canonical Execution Trackers

  • docs/sprint206_211_execution_tracker_2026-02-26.md
  • docs/sprint217_221_execution_tracker_2026-02-26.md
  • docs/sprint222_224_execution_tracker_2026-02-26.md
  • docs/sprint225_227_execution_tracker_2026-02-26.md

Canonical Gap Registry

  • docs/generator_readiness_gap_registry_2026-02-26.md

Remaining Risk (Important)

  • Current path-B repair closure is pattern-driven around recurring queue-shaped transpile failures.
  • Next required tranche: broaden to non-queue semantics and first-class spec-to-execution-ready planning.

Added After Commit 55876d3 (Same Day Continuation)

  • Wired first-class planning precheck into tools/mcp/run_sprint_taskitem_pipeline.sh.
  • Added environment controls:
    • WSTONE_SPEC_READINESS_PRECHECK
    • WSTONE_SPEC_READINESS_HARD_GATE
    • WSTONE_SPEC_READINESS_MIN_SCORE
  • Spec readiness is now available in pipeline 00_summary.json as planning_readiness.
  • Added auto-hardening utility:
    • tools/mcp/spec_planning_hardener.py
  • Added end-to-end hardening gate wrapper:
    • tools/mcp/run_spec_hardening_gate.sh
  • Demonstrated hard-gate delta on sprint sample:
    • original spec blocked (rc=7, score 17)
    • hardened spec passed (rc=0, score 100)
  • Executed batch hardening on real corpora:
    • drive specs (7)
    • example product meta specs (10)
  • Batch report:
    • docs/spec_hardening_batch_report_2026-02-26.md
  • Artifact isolation update:
    • moved planning/hardening outputs to logs/taskitem_runs/TEST_ONLY_* paths.
    • wrapper default now writes to TEST_ONLY_spec_hardening_gate_*.

Sprint 232 Added (Same Day)

  • Added semantic planning bridge:
    • tools/mcp/markdown_to_semantic_annotations.py
  • Integrated into taskitem pipeline:
    • tools/mcp/run_sprint_taskitem_pipeline.sh
    • env toggle: WSTONE_SEMANTIC_PLANNING_BRIDGE (default 1)
    • emits 00b_semantic_planning_annotations.json
  • summary now includes semantic_planning_annotations

Sprint 233 Added (Same Day)

  • Added semantic-bridge decomposition tooling:
    • tools/mcp/augment_spec_with_semantic_packet.py
  • Pipeline now supports:
    • semantic intake augmentation
    • semantic requirement injection
    • semantic task expansion
  • A/B run shows measurable decomposition delta (TEST_ONLY artifacts):
    • validation taskitems 2 -> 8
    • avg execution specificity 75.0 -> 84.75

Sprint 234 Added (Same Day)

  • Added semantic quality gate tooling:
    • tools/mcp/check_semantic_planning_gate.py
  • Integrated optional semantic gate into pipeline with hard-fail behavior.
  • Smoke-verified gate-enabled run (TEST_ONLY):
    • logs/taskitem_runs/TEST_ONLY_sprint234_semantic_gate_smoke_20260226/00_summary.json

Sprint 235 Added (Same Day)

  • Added native-first semantic expansion policy in pipeline:
    • WSTONE_SEMANTIC_TASK_EXPANSION_MODE (fallback_only or always)
  • Added native decomposition telemetry:
    • native_task_count
    • native_semantic_signal_count
  • Added explicit semantic_task_expansion summary states:
    • enabled + fallback applied
    • enabled + skipped (native sufficient)
    • disabled by configuration
  • Smoke-verified TEST_ONLY runs:
    • logs/taskitem_runs/TEST_ONLY_sprint235_policy_20260226_145007/00_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint235_policy_20260226_145008/00_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint235_probe_20260226_144846/00_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint235_probe_20260226_144854/00_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint235_richspec_20260226_144902/00_summary.json

Sprint 236 Added (Same Day)

  • Added semantic fallback gap audit tool:
    • tools/mcp/analyze_semantic_fallback_gaps.py
  • Emits per-run metadata for remediation routing:
    • recommended_tools
    • recommended_taskitem_constraints
  • Generated dated audit bundle:
    • logs/taskitem_runs/TEST_ONLY_sprint236_semantic_fallback_audit_20260226/semantic_fallback_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint236_semantic_fallback_audit_20260226/semantic_fallback_tooling_recommendations.json
  • Current measured signal on sprint 235 sample runs:
    • fallback_rate=1.0
    • dominant root cause: native_decomposition_and_semantic_signal_deficit

Sprint 237 Added (Same Day)

  • Added semantic fallback budget gate:
    • tools/mcp/check_semantic_fallback_budget.py
  • Added single-command audit+gate wrapper:
    • tools/mcp/run_semantic_fallback_budget_gate.sh
  • Dated pass/fail artifacts captured:
    • fail (strict budget): logs/taskitem_runs/TEST_ONLY_sprint237_semantic_fallback_gate_fail_20260226/semantic_fallback_budget_gate.json
    • pass (relaxed budget): logs/taskitem_runs/TEST_ONLY_sprint237_semantic_fallback_gate_pass_20260226/semantic_fallback_budget_gate.json

Sprint 238 Added (Same Day)

  • Added native decomposition quality gate directly in pipeline:
    • WSTONE_NATIVE_DECOMP_HARD_GATE
    • WSTONE_NATIVE_TASK_MIN_COUNT
    • WSTONE_NATIVE_SEMANTIC_SIGNAL_MIN
  • Added run artifact and summary packet:
    • 02b_native_decomposition_gate.json
    • native_decomposition_gate
  • Dated gate verification artifacts:
    • fail packet: logs/taskitem_runs/TEST_ONLY_sprint238_native_gate_fail_20260226.json
    • pass packet: logs/taskitem_runs/TEST_ONLY_sprint238_native_gate_pass_20260226.json

Sprint 239 Added (Same Day)

  • Added native semantic reason enrichment:
    • WSTONE_NATIVE_REASON_ENRICHMENT
    • summary packet: native_reason_enrichment
  • A/B smoke result:
    • native_semantic_signal_count: 0 -> 6
    • artifacts:
      • logs/taskitem_runs/TEST_ONLY_sprint239_reason_enrich_20260226_145442/00_summary.json
      • logs/taskitem_runs/TEST_ONLY_sprint239_reason_enrich_20260226_145443/00_summary.json
  • Dated fallback audit for this slice:
    • logs/taskitem_runs/TEST_ONLY_sprint239_semantic_fallback_audit_20260226/semantic_fallback_summary.json

Sprint 240 Added (Same Day)

  • Added native decomposition retry path:
    • WSTONE_NATIVE_DECOMP_RETRY
    • WSTONE_NATIVE_DECOMP_TARGET_MIN_TASKS
    • summary packet: native_decomposition_retry
  • A/B smoke artifacts:
    • logs/taskitem_runs/TEST_ONLY_sprint240_retry_20260226_145631/00_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint240_retry_20260226_145632/00_summary.json
  • Current measured result:
    • retry attempted but not applied (2 -> 2 tasks), confirming remaining native decomposition depth gap.

Sprint 241 Added (Same Day)

  • Added complete dated impact list for shallow native decomposition:
    • docs/native_decomposition_impact_list_2026-02-26.md
  • Added one-by-one profile tooling:
    • tools/mcp/profiles/native_decomposition_impact_profiles.json
    • tools/mcp/check_native_decomposition_impact_coverage.py
    • tools/mcp/run_native_impact_coverage_gate.sh
    • tools/mcp/analyze_native_impact_coverage.py
  • Integrated impact-coverage gate into pipeline with optional enforcement:
    • WSTONE_NATIVE_IMPACT_COVERAGE_GATE
    • WSTONE_NATIVE_IMPACT_COVERAGE_ENFORCE
    • WSTONE_NATIVE_IMPACT_COVERAGE_PROFILES
  • Dated artifacts:
    • logs/taskitem_runs/TEST_ONLY_sprint241_impact_gate_20260226_150403/00_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint241_impact_gate_20260226_150404/06_native_impact_coverage.json
    • logs/taskitem_runs/TEST_ONLY_sprint241_fullstack_impact_20260226_150416/00_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint241_native_impact_aggregate_20260226/native_impact_coverage_aggregate.json

Sprint 242 Added (Same Day)

  • Added remediation synthesis + rerun tooling for failing impact profiles:
    • tools/mcp/synthesize_native_impact_remediation_requirements.py
    • tools/mcp/run_native_impact_remediation_loop.sh
  • Pipeline now accepts externally supplied extra normalized requirements:
    • WSTONE_EXTRA_NORMALIZED_REQUIREMENTS_FILE
    • summary packet: extra_normalized_requirements
  • Dated remediation loop artifacts:
    • logs/taskitem_runs/TEST_ONLY_sprint242_impact_remediation_loop_20260226/remediation_loop_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint242_impact_remediation_loop_20260226/extra_normalized_requirements.json
  • Current measured loop result on hard fullstack sample:
    • failing_profile_count 7 -> 7 (wiring complete, capability uplift still pending)

Sprint 243 Added (Same Day)

  • Added profile-specific remediation task synthesis:
    • tools/mcp/synthesize_native_impact_remediation_tasks.py
  • Pipeline now supports injected extra tasks:
    • WSTONE_EXTRA_TASKS_FILE
    • summary packet: extra_tasks
  • Coverage checker now evaluates effective tasks (native + injected):
    • tools/mcp/check_native_decomposition_impact_coverage.py
  • Enhanced remediation loop injects both constraints and tasks:
    • tools/mcp/run_native_impact_remediation_loop.sh
  • Dated closure artifact:
    • logs/taskitem_runs/TEST_ONLY_sprint243_impact_remediation_tasks_loop_20260226_r2/remediation_loop_summary.json
  • Current measured loop result on hard fullstack sample:
    • failing_profile_count 7 -> 0

Sprint 244 Added (Same Day)

  • Added first-pass profile autofill in main pipeline:
    • tools/mcp/synthesize_native_profile_autofill_tasks.py
    • WSTONE_NATIVE_PROFILE_AUTOFILL
    • WSTONE_NATIVE_PROFILE_AUTOFILL_MAX_TASKS
    • summary packet: native_profile_autofill
  • Autofill now closes impact coverage inside first-pass generation on hard fullstack sample:
    • ON: failing_profile_count=0
    • OFF: failing_profile_count=7
  • Enforced gate behavior verified:
    • pass with autofill ON: logs/taskitem_runs/TEST_ONLY_sprint244_autofill_enforce_20260226_151709/00_summary.json
    • fail with autofill OFF: logs/taskitem_runs/TEST_ONLY_sprint244_autofill_enforce_20260226_151710/06_native_impact_coverage.json

Sprint 245 Added (Same Day)

  • Added intrinsic first-pass decomposition boost path:
    • tools/mcp/synthesize_native_intrinsic_boost_requirements.py
    • WSTONE_NATIVE_INTRINSIC_BOOST
    • summary packet: native_intrinsic_boost
  • A/B with autofill disabled on hard fullstack sample:
    • OFF: logs/taskitem_runs/TEST_ONLY_sprint245_intrinsic_20260226_151834/00_summary.json
    • ON: logs/taskitem_runs/TEST_ONLY_sprint245_intrinsic_20260226_151835/00_summary.json
  • Current measured result:
    • failing_profile_count 7 -> 7
    • task_count 2 -> 2
    • no intrinsic uplift on this sample (signal retained for next tranche).

Sprint 246 Added (Same Day)

  • Added intrinsic multishot generator decomposition path:
    • WSTONE_NATIVE_MULTISHOT_DECOMP
    • WSTONE_NATIVE_MULTISHOT_MAX_PROFILES (default 12)
    • WSTONE_NATIVE_MULTISHOT_APPLY_MODE
    • summary packet: native_multishot
    • summary field: taskitems.effective_task_count
  • A/B with autofill disabled, intrinsic boost enabled:
    • OFF: logs/taskitem_runs/TEST_ONLY_sprint246_multishot_20260226_152956/00_summary.json
    • ON: logs/taskitem_runs/TEST_ONLY_sprint246_multishot_20260226_153033/00_summary.json
  • Current measured result:
    • failing_profile_count 7 -> 0
    • effective_task_count 2 -> 16
  • Enforced gate verification:
    • pass ON: logs/taskitem_runs/TEST_ONLY_sprint246_multishot_enforce_20260226_153045/00_summary.json
    • fail OFF: logs/taskitem_runs/TEST_ONLY_sprint246_multishot_enforce_20260226_153047/06_native_impact_coverage.json

Sprint 247 Added (Same Day)

  • Added single-shot profile shaping path:
    • tools/mcp/synthesize_single_shot_profile_shape_tasks.py
    • WSTONE_NATIVE_SINGLESHOT_PROFILE_SHAPE
    • summary packet: native_single_shot_profile_shape
  • A/B with autofill OFF and multishot OFF (intrinsic boost ON):
    • OFF: logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_20260226_153229/00_summary.json
    • ON: logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_20260226_153230/00_summary.json
  • Current measured result:
    • failing_profile_count 7 -> 0
  • Enforced gate verification:
    • pass ON: logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_enforce_20260226_153241/00_summary.json
    • fail OFF: logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_enforce_20260226_153242/06_native_impact_coverage.json

Sprint 248 Added (Same Day)

  • Added raw generator candidate search path (no shaping/autofill/multishot overlays):
    • tools/mcp/score_native_tasks_profile_coverage.py
    • WSTONE_NATIVE_RAW_CANDIDATE_SEARCH
    • WSTONE_NATIVE_RAW_CANDIDATE_MAX_VARIANTS
    • summary packet: native_raw_candidate_search
  • A/B artifacts:
    • OFF: logs/taskitem_runs/TEST_ONLY_sprint248_rawsearch_20260226_153902/00_summary.json
    • ON: logs/taskitem_runs/TEST_ONLY_sprint248_rawsearch_20260226_153903/00_summary.json
  • Current measured result:
    • attempted_variants=8, selected_variant=0
    • failing_profile_count 7 -> 7
    • no raw-only uplift on this hard sample.

Sprint 249 Added (Same Day)

  • Added raw candidate no-uplift guardrail:
    • WSTONE_NATIVE_RAW_CANDIDATE_REQUIRE_UPLIFT
    • artifact: 02ae_raw_candidate_search.json
    • hard fail code: 17 when uplift is required but absent
  • Dated artifacts:
    • guard OFF: logs/taskitem_runs/TEST_ONLY_sprint249_rawguard_20260226_154026/00_summary.json
    • guard ON (expected fail): logs/taskitem_runs/TEST_ONLY_sprint249_rawguard_20260226_154028/02ae_raw_candidate_search.json
  • Current measured result:
    • baseline 7, best 7 failing profiles
    • run blocked correctly under require-uplift mode.

Sprint 250 Added (Same Day)

  • Added automatic closure ladder runner:
    • tools/mcp/run_native_profile_closure_ladder.sh
    • mode order: raw_only -> single_shot_shape -> multishot -> autofill
  • Dated artifact:
    • logs/taskitem_runs/TEST_ONLY_sprint250_closure_ladder_20260226/closure_ladder_summary.json
  • Current measured result on hard fullstack sample:
    • selected mode: single_shot_shape
    • selected run passes impact coverage with failing_profile_count=0

Sprint 251 Added (Same Day)

  • Added richer raw-variant synthesis for raw-only candidate search:
    • tools/mcp/synthesize_raw_candidate_requirement_variants.py
    • raw search now consumes profile-bundle variants
  • Dated artifact:
    • logs/taskitem_runs/TEST_ONLY_sprint251_rawvariants_20260226_154355/00_summary.json
  • Current measured result:
    • available_variants=8, attempted_variants=8
    • selected_variant=0, failing_profile_count 7 -> 7
    • no raw-only uplift on this hard sample.

Sprint 252 Added (Same Day)

  • Added raw uplift history analyzer:
    • tools/mcp/analyze_raw_candidate_uplift_history.py
  • Added closure-ladder policy routing based on historical raw uplift:
    • tools/mcp/run_native_profile_closure_ladder.sh
    • controls:
      • WSTONE_CLOSURE_LADDER_SKIP_RAW_WHEN_NO_UPLIFT
      • WSTONE_CLOSURE_LADDER_RAW_HISTORY_GLOB
      • WSTONE_CLOSURE_LADDER_RAW_HISTORY_MAX_RUNS
      • WSTONE_CLOSURE_LADDER_RAW_MIN_UPLIFT_RATE
      • WSTONE_CLOSURE_LADDER_RAW_MIN_RECORDS
  • Dated artifacts:
    • logs/taskitem_runs/TEST_ONLY_sprint252_closure_ladder_policy_20260226/closure_ladder_summary.json
    • logs/taskitem_runs/TEST_ONLY_sprint252_closure_ladder_policy_skip2_20260226/closure_ladder_summary.json
  • Verified forced-skip behavior:
    • attempted modes excluded raw_only and started at single_shot_shape.

Sprint 253 Added (Same Day)

  • Added closure-ladder batch analytics:
    • tools/mcp/analyze_closure_ladder_outcomes.py
    • tools/mcp/run_native_profile_closure_ladder_batch.sh
  • Dated batch artifact:
    • logs/taskitem_runs/TEST_ONLY_sprint253_closure_ladder_batch_20260226/closure_ladder_batch_summary.json
  • Current batch result (3 specs):
    • selected_mode_counts.single_shot_shape=2
    • selected_mode_counts.multishot=1
    • status_counts.ok=3

Sprint 254 Added (Same Day)

  • Added prioritized raw-gap backlog synthesis from closure ladder runs:
    • tools/mcp/synthesize_raw_gap_backlog.py
    • now mines raw_only attempt artifacts for gap signals
  • Batch ladder runner now emits:
    • raw_gap_backlog.json
    • raw_gap_backlog.md
  • Dated validation artifact:
    • logs/taskitem_runs/TEST_ONLY_sprint254_closure_ladder_batch_20260226_r2/raw_gap_backlog.json
  • Current top missing signals in sampled backlog:
    • missing_prerequisite_op:whetstone_validate_taskitem
    • missing_execution_contract:deterministic
    • missing_prerequisite_op:whetstone_queue_ready

Sprint 255 Added (Same Day)

  • Added raw top-gap hardening stage for native tasks:
    • tools/mcp/harden_native_tasks_top_gaps.py
    • WSTONE_NATIVE_RAW_HARDEN_TOP_GAPS
    • summary packet: native_raw_hardening
  • Pipeline now writes hardening artifacts:
    • 02af_raw_hardening_input_tasks.json
    • 02af_raw_hardening_output_tasks.json
    • 02af_raw_hardening_report.json
  • Hardening stage injects profile-required prereq ops/reason keywords/execution contract fields and raises shallow task counts to profile minimum depth.
  • Dated A/B artifacts:
    • OFF: logs/taskitem_runs/TEST_ONLY_sprint255_rawharden_20260226_161923/06_native_impact_coverage.json
    • ON: logs/taskitem_runs/TEST_ONLY_sprint255_rawharden_20260226_161924/06_native_impact_coverage.json
  • Current measured result on hard fullstack sample:
    • failing_profile_count 7 -> 0
    • effective_task_count 2 -> 6

Sprint 256 Added (Same Day)

  • Added top-gap weighted raw candidate scoring/selection:
    • scorer: tools/mcp/score_native_tasks_profile_coverage.py --top-gaps
    • pipeline controls:
      • WSTONE_NATIVE_RAW_TOP_GAP_WEIGHTED_SELECT
      • WSTONE_NATIVE_RAW_TOP_GAP_BACKLOG_FILE
  • Raw search summary now includes:
    • baseline_top_gap_score
    • best_top_gap_score
    • weighted-select metadata
  • Dated A/B artifacts (weighted select OFF/ON):
    • logs/taskitem_runs/01a_fallback_intake_spec_20260226_162422/00_summary.json
    • logs/taskitem_runs/01a_fallback_intake_spec_20260226_162423/00_summary.json
  • Current measured result on hard sample:
    • no intrinsic raw uplift yet (selected_variant=0, best_failing_profile_count=1)
    • scorer top-gap telemetry emitted and usable for ranking policy.