From fd8d582db774a38062b31e73fdd6dbf8b466578f Mon Sep 17 00:00:00 2001 From: Bill Date: Thu, 26 Feb 2026 16:27:39 -0700 Subject: [PATCH] Add top-gap uplift hard gate for raw candidate search (sprint 257) --- docs/constructive_editing_runtime_plan.md | 3 ++ ...rator_readiness_gap_registry_2026-02-26.md | 2 +- docs/progress_log_2026-02-26.md | 14 ++++++++ .../sprint257_execution_tracker_2026-02-26.md | 33 +++++++++++++++++++ editor/src/Sprint257IntegrationSummary.h | 6 ++++ sprint257_plan.md | 11 +++++++ tools/mcp/run_sprint_taskitem_pipeline.sh | 16 ++++++++- 7 files changed, 83 insertions(+), 2 deletions(-) create mode 100644 docs/sprint257_execution_tracker_2026-02-26.md create mode 100644 editor/src/Sprint257IntegrationSummary.h create mode 100644 sprint257_plan.md diff --git a/docs/constructive_editing_runtime_plan.md b/docs/constructive_editing_runtime_plan.md index 9c90f87..afb5e6c 100644 --- a/docs/constructive_editing_runtime_plan.md +++ b/docs/constructive_editing_runtime_plan.md @@ -220,3 +220,6 @@ Planning runtime controls (active): - `WSTONE_NATIVE_RAW_TOP_GAP_WEIGHTED_SELECT` - `WSTONE_NATIVE_RAW_TOP_GAP_BACKLOG_FILE` - scorer `--top-gaps` telemetry (`top_gap_score`) +- raw candidate search can enforce top-gap uplift as hard policy: + - `WSTONE_NATIVE_RAW_TOP_GAP_REQUIRE_UPLIFT` + - failure code `18` when no weighted uplift is achieved diff --git a/docs/generator_readiness_gap_registry_2026-02-26.md b/docs/generator_readiness_gap_registry_2026-02-26.md index f012a90..4d1fc43 100644 --- a/docs/generator_readiness_gap_registry_2026-02-26.md +++ b/docs/generator_readiness_gap_registry_2026-02-26.md @@ -47,7 +47,7 @@ This is the canonical dated registry for "not production-ready" generator gaps. | GR-018 | Performance/security/rollout constrained refactor enforcement gap | `docs/gap_hunt_fullstack_multifile_2026-02-26.md`, `logs/taskitem_runs/challenging_fullstack_multifile_20260226_r5/results.jsonl` | `partial` | Hard checks now enforce migration rollback + data-loss policy, security deny-by-default, SLO p95 presence, and rollout staged+abort policy. Enforcement is contract-level; generator capability under these constraints is still weak in C++ AB path. | Constraint-aware generation follow-up | | GR-019 | Parity-blocked readiness load (gating without capability closure) | `logs/taskitem_runs/challenging_fullstack_multifile_20260226_r7/summary.json`, `logs/taskitem_runs/challenging_subset_prod_20260226_r7/summary.json`, `docs/sprint225_227_execution_tracker_2026-02-26.md` | `partial` | Sprint 225-227 reduced blocked parity load from `6` to `0` on both tracked hard catalogs while keeping unresolved divergence at `0`. This closes immediate safety debt for current corpora, but robustness is still contingent on pattern-driven repair classes. | Generalize repairs beyond queue-shaped transpile outputs | | GR-020 | Semantic fallback overuse masks weak native decomposition | `logs/taskitem_runs/TEST_ONLY_sprint236_semantic_fallback_audit_20260226/semantic_fallback_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint237_semantic_fallback_gate_fail_20260226/semantic_fallback_budget_gate.json`, `logs/taskitem_runs/TEST_ONLY_sprint238_native_gate_fail_20260226.json`, `logs/taskitem_runs/TEST_ONLY_sprint239_semantic_fallback_audit_20260226/semantic_fallback_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint240_retry_20260226_145631/00_summary.json`, `docs/sprint236_execution_tracker_2026-02-26.md`, `docs/sprint237_execution_tracker_2026-02-26.md`, `docs/sprint238_execution_tracker_2026-02-26.md`, `docs/sprint239_execution_tracker_2026-02-26.md`, `docs/sprint240_execution_tracker_2026-02-26.md` | `partial` | Sprint 236 added deterministic fallback-gap auditing with tool + constraint metadata. Sprint 237 added fallback-budget hard gate (`max-fallback-rate`) and produced expected fail/pass artifacts. Sprint 238 added native decomposition hard gate in pipeline (`min task count`, `min semantic signal count`) so weak native generation can be blocked before fallback masking. Sprint 239 added native reason enrichment and improved semantic signal density (`0 -> 6`). Sprint 240 added native decomposition retry with explicit minimum-task policy; on current sample retry attempted but did not improve depth (`2 -> 2`). | Native decomposition task-depth upgrade (increase native task granularity beyond 2) | -| GR-021 | Impact-specific native decomposition coverage not enforced uniformly | `docs/native_decomposition_impact_list_2026-02-26.md`, `tools/mcp/profiles/native_decomposition_impact_profiles.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_fullstack_impact_20260226_150416/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_native_impact_aggregate_20260226/native_impact_coverage_aggregate.json`, `logs/taskitem_runs/TEST_ONLY_sprint242_impact_remediation_loop_20260226/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint243_impact_remediation_tasks_loop_20260226_r2/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_20260226_151651/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_enforce_20260226_151709/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint245_intrinsic_20260226_151835/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_20260226_153033/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_enforce_20260226_153045/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_20260226_153230/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_enforce_20260226_153241/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint248_rawsearch_20260226_153903/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint249_rawguard_20260226_154028/02ae_raw_candidate_search.json`, `logs/taskitem_runs/TEST_ONLY_sprint250_closure_ladder_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint251_rawvariants_20260226_154355/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint252_closure_ladder_policy_skip2_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint253_closure_ladder_batch_20260226/closure_ladder_batch_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint254_closure_ladder_batch_20260226_r2/raw_gap_backlog.json`, `logs/taskitem_runs/TEST_ONLY_sprint255_rawharden_20260226_161924/00_summary.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162423/00_summary.json` | `partial` | Sprint 241 introduced impact profiles/gates; 242-243 remediation loops; 244 first-pass autofill; 245 intrinsic constraints only (no uplift); 246 multishot closure; 247 single-shot shaping closure; 248 raw candidate search no uplift; 249 no-uplift guardrail; 250 closure ladder; 251 richer raw variants no uplift; 252 history-aware ladder routing; 253 batch ladder analytics; 254 raw gap backlog synthesis from `raw_only` attempts identifies top missing raw signals; 255 adds first-pass raw top-gap hardening that closes hard-sample profile coverage (`failing_profile_count 7 -> 0`) by injecting missing ops/contracts/reasons and minimum task depth; 256 adds top-gap weighted candidate selection telemetry and tie-break policy, but measured sample still shows no intrinsic raw uplift (`selected_variant=0`). Residual gap remains in intrinsic raw generator quality because coverage is still recovered primarily via post-generation hardening/shaping paths. | Add intrinsic raw generation hard gate requiring minimum top-gap score uplift (or fail fast) before promoting raw-only path | +| GR-021 | Impact-specific native decomposition coverage not enforced uniformly | `docs/native_decomposition_impact_list_2026-02-26.md`, `tools/mcp/profiles/native_decomposition_impact_profiles.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_fullstack_impact_20260226_150416/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_native_impact_aggregate_20260226/native_impact_coverage_aggregate.json`, `logs/taskitem_runs/TEST_ONLY_sprint242_impact_remediation_loop_20260226/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint243_impact_remediation_tasks_loop_20260226_r2/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_20260226_151651/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_enforce_20260226_151709/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint245_intrinsic_20260226_151835/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_20260226_153033/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_enforce_20260226_153045/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_20260226_153230/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_enforce_20260226_153241/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint248_rawsearch_20260226_153903/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint249_rawguard_20260226_154028/02ae_raw_candidate_search.json`, `logs/taskitem_runs/TEST_ONLY_sprint250_closure_ladder_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint251_rawvariants_20260226_154355/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint252_closure_ladder_policy_skip2_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint253_closure_ladder_batch_20260226/closure_ladder_batch_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint254_closure_ladder_batch_20260226_r2/raw_gap_backlog.json`, `logs/taskitem_runs/TEST_ONLY_sprint255_rawharden_20260226_161924/00_summary.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162423/00_summary.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162633/02ae_raw_candidate_search.json` | `partial` | Sprint 241 introduced impact profiles/gates; 242-243 remediation loops; 244 first-pass autofill; 245 intrinsic constraints only (no uplift); 246 multishot closure; 247 single-shot shaping closure; 248 raw candidate search no uplift; 249 no-uplift guardrail; 250 closure ladder; 251 richer raw variants no uplift; 252 history-aware ladder routing; 253 batch ladder analytics; 254 raw gap backlog synthesis from `raw_only` attempts identifies top missing raw signals; 255 adds first-pass raw top-gap hardening that closes hard-sample profile coverage (`failing_profile_count 7 -> 0`) by injecting missing ops/contracts/reasons and minimum task depth; 256 adds top-gap weighted candidate selection telemetry and tie-break policy; 257 adds hard fail policy when top-gap uplift is required but absent (`exit 18`). Residual gap remains intrinsic: raw generation still fails to produce uplift on this sample without hardening/shaping overlays. | Add intrinsic raw requirement synthesis loop that explicitly targets top weighted missing signals before raw generation retries | ## What Was Covered Today (Sprints 175-184) diff --git a/docs/progress_log_2026-02-26.md b/docs/progress_log_2026-02-26.md index 59230ec..f9834a2 100644 --- a/docs/progress_log_2026-02-26.md +++ b/docs/progress_log_2026-02-26.md @@ -427,3 +427,17 @@ - Current measured result on hard sample: - no intrinsic raw uplift yet (`selected_variant=0`, `best_failing_profile_count=1`) - scorer top-gap telemetry emitted and usable for ranking policy. + +## Sprint 257 Added (Same Day) + +- Added top-gap uplift hard gate in raw candidate search: + - `WSTONE_NATIVE_RAW_TOP_GAP_REQUIRE_UPLIFT` + - exit code `18` on missing weighted mode config or no top-gap uplift +- Raw search summary now includes: + - `top_gap_require_uplift` +- Dated validation artifacts: + - pass (gate OFF): `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162632/00_summary.json` + - fail (gate ON): `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162633/02ae_raw_candidate_search.json` +- Current measured result on hard sample: + - `baseline_top_gap_score=0`, `best_top_gap_score=0` + - run blocked correctly when uplift requirement is enabled. diff --git a/docs/sprint257_execution_tracker_2026-02-26.md b/docs/sprint257_execution_tracker_2026-02-26.md new file mode 100644 index 0000000..72fc055 --- /dev/null +++ b/docs/sprint257_execution_tracker_2026-02-26.md @@ -0,0 +1,33 @@ +# Sprint 257 Execution Tracker - 2026-02-26 + +## Scope +- `sprint257_plan.md` + +## Implemented + +Pipeline gating enhancement in `tools/mcp/run_sprint_taskitem_pipeline.sh`: +- Added: + - `WSTONE_NATIVE_RAW_TOP_GAP_REQUIRE_UPLIFT` +- Raw candidate search now hard-fails with exit code `18` when: + - uplift is required but weighted top-gap mode is not configured (`WSTONE_NATIVE_RAW_TOP_GAP_WEIGHTED_SELECT=1` + valid backlog file), or + - `best_top_gap_score <= baseline_top_gap_score` +- Raw search summary now emits: + - `top_gap_require_uplift` + +## Validation Artifacts + +Gate OFF (expected pass): +- `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162632/00_summary.json` + +Gate ON (expected fail): +- `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162633/02ae_raw_candidate_search.json` +- command exit code: `18` + +Observed result on hard sample: +- `baseline_top_gap_score=0` +- `best_top_gap_score=0` +- with required uplift enabled, run blocked as designed. + +## Explicit Completion Signal + +- Sprint 257: `DONE` (implemented + enforced no-uplift guard for top-gap policy) diff --git a/editor/src/Sprint257IntegrationSummary.h b/editor/src/Sprint257IntegrationSummary.h new file mode 100644 index 0000000..d8e51c5 --- /dev/null +++ b/editor/src/Sprint257IntegrationSummary.h @@ -0,0 +1,6 @@ +#pragma once + +// Sprint 257 integration summary: +// - Added hard gate requiring top-gap score uplift for raw candidate search. +// - Gate returns exit code 18 when weighted top-gap mode is misconfigured or no uplift is observed. +// - Raw search summary now includes top-gap uplift policy metadata. diff --git a/sprint257_plan.md b/sprint257_plan.md new file mode 100644 index 0000000..4044db7 --- /dev/null +++ b/sprint257_plan.md @@ -0,0 +1,11 @@ +# Sprint 257 Plan: Top-Gap Uplift Hard Gate + +## Goal +Convert top-gap weighted raw scoring from advisory telemetry into an enforceable policy gate so non-improving raw-only runs fail fast. + +## Steps +- Step 2332: Add raw top-gap uplift gate env control. +- Step 2333: Fail raw search when weighted top-gap mode is not properly configured. +- Step 2334: Fail raw search when best top-gap score does not improve baseline. +- Step 2335: Emit top-gap uplift gate policy fields in raw search summary. +- Step 2336: Add `Sprint257IntegrationSummary.h` and execution tracker. diff --git a/tools/mcp/run_sprint_taskitem_pipeline.sh b/tools/mcp/run_sprint_taskitem_pipeline.sh index c38e3a5..da78d4f 100755 --- a/tools/mcp/run_sprint_taskitem_pipeline.sh +++ b/tools/mcp/run_sprint_taskitem_pipeline.sh @@ -55,6 +55,7 @@ NATIVE_RAW_CANDIDATE_REQUIRE_UPLIFT="${WSTONE_NATIVE_RAW_CANDIDATE_REQUIRE_UPLIF NATIVE_RAW_HARDEN_TOP_GAPS="${WSTONE_NATIVE_RAW_HARDEN_TOP_GAPS:-0}" NATIVE_RAW_TOP_GAP_WEIGHTED_SELECT="${WSTONE_NATIVE_RAW_TOP_GAP_WEIGHTED_SELECT:-0}" NATIVE_RAW_TOP_GAP_BACKLOG_FILE="${WSTONE_NATIVE_RAW_TOP_GAP_BACKLOG_FILE:-}" +NATIVE_RAW_TOP_GAP_REQUIRE_UPLIFT="${WSTONE_NATIVE_RAW_TOP_GAP_REQUIRE_UPLIFT:-0}" EXTRA_NORMALIZED_REQUIREMENTS_FILE="${WSTONE_EXTRA_NORMALIZED_REQUIREMENTS_FILE:-}" EXTRA_TASKS_FILE="${WSTONE_EXTRA_TASKS_FILE:-}" CAPABILITY_SIGNALS_JSON="${WSTONE_CAPABILITY_SIGNALS_JSON:-}" @@ -489,13 +490,26 @@ if [[ "$NATIVE_RAW_CANDIDATE_SEARCH" == "1" ]]; then --argjson available_variants "$variant_total" \ --argjson top_gap_weighted_select "$use_top_gap_weighted_select" \ --arg top_gap_backlog_file "$NATIVE_RAW_TOP_GAP_BACKLOG_FILE" \ - '{enabled:$enabled, attempted_variants:$attempted, successful_variants:$successful, available_variants:$available_variants, selected_variant:$best_variant, baseline_failing_profile_count:$baseline_failing_profile_count, baseline_task_count:$baseline_task_count, baseline_top_gap_score:$baseline_top_gap_score, best_failing_profile_count:$best_failing_profile_count, best_task_count:$best_task_count, best_top_gap_score:$best_top_gap_score, top_gap_weighted_select:$top_gap_weighted_select, top_gap_backlog_file:$top_gap_backlog_file}')" + --argjson top_gap_require_uplift "$NATIVE_RAW_TOP_GAP_REQUIRE_UPLIFT" \ + '{enabled:$enabled, attempted_variants:$attempted, successful_variants:$successful, available_variants:$available_variants, selected_variant:$best_variant, baseline_failing_profile_count:$baseline_failing_profile_count, baseline_task_count:$baseline_task_count, baseline_top_gap_score:$baseline_top_gap_score, best_failing_profile_count:$best_failing_profile_count, best_task_count:$best_task_count, best_top_gap_score:$best_top_gap_score, top_gap_weighted_select:$top_gap_weighted_select, top_gap_backlog_file:$top_gap_backlog_file, top_gap_require_uplift:$top_gap_require_uplift}')" printf '%s\n' "$NATIVE_RAW_CANDIDATE_SEARCH_JSON" > "$OUT_DIR/02ae_raw_candidate_search.json" if [[ "$NATIVE_RAW_CANDIDATE_REQUIRE_UPLIFT" == "1" && "$best_fail" -ge "$baseline_fail" ]]; then echo "error: raw candidate search did not improve failing profile count" >&2 echo "error: see $OUT_DIR/02ae_raw_candidate_search.json" >&2 exit 17 fi + if [[ "$NATIVE_RAW_TOP_GAP_REQUIRE_UPLIFT" == "1" ]]; then + if [[ "$use_top_gap_weighted_select" != true ]]; then + echo "error: raw top-gap uplift requires weighted select with a valid backlog file" >&2 + echo "error: set WSTONE_NATIVE_RAW_TOP_GAP_WEIGHTED_SELECT=1 and WSTONE_NATIVE_RAW_TOP_GAP_BACKLOG_FILE=" >&2 + exit 18 + fi + if [[ "$best_top_gap_score" -le "$baseline_top_gap_score" ]]; then + echo "error: raw candidate search did not improve top-gap score" >&2 + echo "error: see $OUT_DIR/02ae_raw_candidate_search.json" >&2 + exit 18 + fi + fi else NATIVE_RAW_CANDIDATE_SEARCH_JSON='{"enabled":false}' fi