Add top-gap weighted raw candidate selection telemetry (sprint 256)

This commit is contained in:
Bill
2026-02-26 16:25:57 -07:00
parent 057fcd4087
commit 39575cb383
8 changed files with 204 additions and 5 deletions

View File

@@ -216,3 +216,7 @@ Planning runtime controls (active):
- `tools/mcp/harden_native_tasks_top_gaps.py`
- `WSTONE_NATIVE_RAW_HARDEN_TOP_GAPS`
- `native_raw_hardening`
- raw candidate selection can be weighted by historical gap priorities:
- `WSTONE_NATIVE_RAW_TOP_GAP_WEIGHTED_SELECT`
- `WSTONE_NATIVE_RAW_TOP_GAP_BACKLOG_FILE`
- scorer `--top-gaps` telemetry (`top_gap_score`)

View File

@@ -47,7 +47,7 @@ This is the canonical dated registry for "not production-ready" generator gaps.
| GR-018 | Performance/security/rollout constrained refactor enforcement gap | `docs/gap_hunt_fullstack_multifile_2026-02-26.md`, `logs/taskitem_runs/challenging_fullstack_multifile_20260226_r5/results.jsonl` | `partial` | Hard checks now enforce migration rollback + data-loss policy, security deny-by-default, SLO p95 presence, and rollout staged+abort policy. Enforcement is contract-level; generator capability under these constraints is still weak in C++ AB path. | Constraint-aware generation follow-up |
| GR-019 | Parity-blocked readiness load (gating without capability closure) | `logs/taskitem_runs/challenging_fullstack_multifile_20260226_r7/summary.json`, `logs/taskitem_runs/challenging_subset_prod_20260226_r7/summary.json`, `docs/sprint225_227_execution_tracker_2026-02-26.md` | `partial` | Sprint 225-227 reduced blocked parity load from `6` to `0` on both tracked hard catalogs while keeping unresolved divergence at `0`. This closes immediate safety debt for current corpora, but robustness is still contingent on pattern-driven repair classes. | Generalize repairs beyond queue-shaped transpile outputs |
| GR-020 | Semantic fallback overuse masks weak native decomposition | `logs/taskitem_runs/TEST_ONLY_sprint236_semantic_fallback_audit_20260226/semantic_fallback_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint237_semantic_fallback_gate_fail_20260226/semantic_fallback_budget_gate.json`, `logs/taskitem_runs/TEST_ONLY_sprint238_native_gate_fail_20260226.json`, `logs/taskitem_runs/TEST_ONLY_sprint239_semantic_fallback_audit_20260226/semantic_fallback_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint240_retry_20260226_145631/00_summary.json`, `docs/sprint236_execution_tracker_2026-02-26.md`, `docs/sprint237_execution_tracker_2026-02-26.md`, `docs/sprint238_execution_tracker_2026-02-26.md`, `docs/sprint239_execution_tracker_2026-02-26.md`, `docs/sprint240_execution_tracker_2026-02-26.md` | `partial` | Sprint 236 added deterministic fallback-gap auditing with tool + constraint metadata. Sprint 237 added fallback-budget hard gate (`max-fallback-rate`) and produced expected fail/pass artifacts. Sprint 238 added native decomposition hard gate in pipeline (`min task count`, `min semantic signal count`) so weak native generation can be blocked before fallback masking. Sprint 239 added native reason enrichment and improved semantic signal density (`0 -> 6`). Sprint 240 added native decomposition retry with explicit minimum-task policy; on current sample retry attempted but did not improve depth (`2 -> 2`). | Native decomposition task-depth upgrade (increase native task granularity beyond 2) |
| GR-021 | Impact-specific native decomposition coverage not enforced uniformly | `docs/native_decomposition_impact_list_2026-02-26.md`, `tools/mcp/profiles/native_decomposition_impact_profiles.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_fullstack_impact_20260226_150416/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_native_impact_aggregate_20260226/native_impact_coverage_aggregate.json`, `logs/taskitem_runs/TEST_ONLY_sprint242_impact_remediation_loop_20260226/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint243_impact_remediation_tasks_loop_20260226_r2/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_20260226_151651/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_enforce_20260226_151709/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint245_intrinsic_20260226_151835/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_20260226_153033/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_enforce_20260226_153045/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_20260226_153230/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_enforce_20260226_153241/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint248_rawsearch_20260226_153903/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint249_rawguard_20260226_154028/02ae_raw_candidate_search.json`, `logs/taskitem_runs/TEST_ONLY_sprint250_closure_ladder_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint251_rawvariants_20260226_154355/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint252_closure_ladder_policy_skip2_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint253_closure_ladder_batch_20260226/closure_ladder_batch_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint254_closure_ladder_batch_20260226_r2/raw_gap_backlog.json`, `logs/taskitem_runs/TEST_ONLY_sprint255_rawharden_20260226_161924/00_summary.json` | `partial` | Sprint 241 introduced impact profiles/gates; 242-243 remediation loops; 244 first-pass autofill; 245 intrinsic constraints only (no uplift); 246 multishot closure; 247 single-shot shaping closure; 248 raw candidate search no uplift; 249 no-uplift guardrail; 250 closure ladder; 251 richer raw variants no uplift; 252 history-aware ladder routing; 253 batch ladder analytics; 254 raw gap backlog synthesis from `raw_only` attempts identifies top missing raw signals; 255 adds first-pass raw top-gap hardening that closes hard-sample profile coverage (`failing_profile_count 7 -> 0`) by injecting missing ops/contracts/reasons and minimum task depth. Residual gap remains in intrinsic raw generator quality because hardening is compensatory post-processing. | Improve intrinsic raw candidate selection/generation so backlog signal closure is achieved without post-hardening overlays |
| GR-021 | Impact-specific native decomposition coverage not enforced uniformly | `docs/native_decomposition_impact_list_2026-02-26.md`, `tools/mcp/profiles/native_decomposition_impact_profiles.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_fullstack_impact_20260226_150416/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_native_impact_aggregate_20260226/native_impact_coverage_aggregate.json`, `logs/taskitem_runs/TEST_ONLY_sprint242_impact_remediation_loop_20260226/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint243_impact_remediation_tasks_loop_20260226_r2/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_20260226_151651/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_enforce_20260226_151709/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint245_intrinsic_20260226_151835/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_20260226_153033/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_enforce_20260226_153045/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_20260226_153230/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_enforce_20260226_153241/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint248_rawsearch_20260226_153903/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint249_rawguard_20260226_154028/02ae_raw_candidate_search.json`, `logs/taskitem_runs/TEST_ONLY_sprint250_closure_ladder_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint251_rawvariants_20260226_154355/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint252_closure_ladder_policy_skip2_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint253_closure_ladder_batch_20260226/closure_ladder_batch_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint254_closure_ladder_batch_20260226_r2/raw_gap_backlog.json`, `logs/taskitem_runs/TEST_ONLY_sprint255_rawharden_20260226_161924/00_summary.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162423/00_summary.json` | `partial` | Sprint 241 introduced impact profiles/gates; 242-243 remediation loops; 244 first-pass autofill; 245 intrinsic constraints only (no uplift); 246 multishot closure; 247 single-shot shaping closure; 248 raw candidate search no uplift; 249 no-uplift guardrail; 250 closure ladder; 251 richer raw variants no uplift; 252 history-aware ladder routing; 253 batch ladder analytics; 254 raw gap backlog synthesis from `raw_only` attempts identifies top missing raw signals; 255 adds first-pass raw top-gap hardening that closes hard-sample profile coverage (`failing_profile_count 7 -> 0`) by injecting missing ops/contracts/reasons and minimum task depth; 256 adds top-gap weighted candidate selection telemetry and tie-break policy, but measured sample still shows no intrinsic raw uplift (`selected_variant=0`). Residual gap remains in intrinsic raw generator quality because coverage is still recovered primarily via post-generation hardening/shaping paths. | Add intrinsic raw generation hard gate requiring minimum top-gap score uplift (or fail fast) before promoting raw-only path |
## What Was Covered Today (Sprints 175-184)

View File

@@ -409,3 +409,21 @@
- Current measured result on hard fullstack sample:
- `failing_profile_count 7 -> 0`
- `effective_task_count 2 -> 6`
## Sprint 256 Added (Same Day)
- Added top-gap weighted raw candidate scoring/selection:
- scorer: `tools/mcp/score_native_tasks_profile_coverage.py --top-gaps`
- pipeline controls:
- `WSTONE_NATIVE_RAW_TOP_GAP_WEIGHTED_SELECT`
- `WSTONE_NATIVE_RAW_TOP_GAP_BACKLOG_FILE`
- Raw search summary now includes:
- `baseline_top_gap_score`
- `best_top_gap_score`
- weighted-select metadata
- Dated A/B artifacts (weighted select OFF/ON):
- `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162422/00_summary.json`
- `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162423/00_summary.json`
- Current measured result on hard sample:
- no intrinsic raw uplift yet (`selected_variant=0`, `best_failing_profile_count=1`)
- scorer top-gap telemetry emitted and usable for ranking policy.

View File

@@ -0,0 +1,53 @@
# Sprint 256 Execution Tracker - 2026-02-26
## Scope
- `sprint256_plan.md`
## Implemented
Scorer enhancement:
- `tools/mcp/score_native_tasks_profile_coverage.py`
- Added optional `--top-gaps <raw_gap_backlog.json>`.
- Computes:
- `top_gap_signal_count`
- `top_gap_signal_hits`
- `top_gap_signal_weight_total`
- `top_gap_score`
Pipeline integration in `tools/mcp/run_sprint_taskitem_pipeline.sh`:
- Added:
- `WSTONE_NATIVE_RAW_TOP_GAP_WEIGHTED_SELECT`
- `WSTONE_NATIVE_RAW_TOP_GAP_BACKLOG_FILE`
- Raw search summary now emits:
- `baseline_top_gap_score`
- `best_top_gap_score`
- `top_gap_weighted_select`
- `top_gap_backlog_file`
- Candidate selection policy:
- first minimize `failing_profile_count`
- if tied and weighted mode enabled, maximize `top_gap_score`
- then maximize `task_count`
## Validation Artifacts
Raw search A/B (weighted selector off/on):
- OFF:
- `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162422/00_summary.json`
- ON:
- `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162423/00_summary.json`
Observed result on hard sample:
- `selected_variant` unchanged (`0`)
- `best_failing_profile_count` unchanged (`1`)
- `best_top_gap_score` remained `0` for this sample
Scorer sanity check on hardened tasks:
- `python3 tools/mcp/score_native_tasks_profile_coverage.py ... --top-gaps ...`
- output: `/tmp/sprint256_score_check.json`
- observed:
- `top_gap_signal_hits=18/20`
- `top_gap_score=86`
## Explicit Completion Signal
- Sprint 256: `DONE` (implemented + measured; no intrinsic uplift yet on this hard sample)