Add raw score/gate parity packet and enforcement (sprint 265)
This commit is contained in:
@@ -253,3 +253,6 @@ Planning runtime controls (active):
|
|||||||
- `native_raw_structural_projector`
|
- `native_raw_structural_projector`
|
||||||
- raw candidate selection now persists selected tasks into generation artifacts for gate parity:
|
- raw candidate selection now persists selected tasks into generation artifacts for gate parity:
|
||||||
- keeps `native_raw_candidate_search` and `native_impact_coverage` on the same task set
|
- keeps `native_raw_candidate_search` and `native_impact_coverage` on the same task set
|
||||||
|
- raw score/gate parity can be emitted and hard-enforced:
|
||||||
|
- `native_raw_score_gate_parity`
|
||||||
|
- `WSTONE_NATIVE_RAW_SCORE_GATE_PARITY_ENFORCE`
|
||||||
|
|||||||
@@ -47,7 +47,7 @@ This is the canonical dated registry for "not production-ready" generator gaps.
|
|||||||
| GR-018 | Performance/security/rollout constrained refactor enforcement gap | `docs/gap_hunt_fullstack_multifile_2026-02-26.md`, `logs/taskitem_runs/challenging_fullstack_multifile_20260226_r5/results.jsonl` | `partial` | Hard checks now enforce migration rollback + data-loss policy, security deny-by-default, SLO p95 presence, and rollout staged+abort policy. Enforcement is contract-level; generator capability under these constraints is still weak in C++ AB path. | Constraint-aware generation follow-up |
|
| GR-018 | Performance/security/rollout constrained refactor enforcement gap | `docs/gap_hunt_fullstack_multifile_2026-02-26.md`, `logs/taskitem_runs/challenging_fullstack_multifile_20260226_r5/results.jsonl` | `partial` | Hard checks now enforce migration rollback + data-loss policy, security deny-by-default, SLO p95 presence, and rollout staged+abort policy. Enforcement is contract-level; generator capability under these constraints is still weak in C++ AB path. | Constraint-aware generation follow-up |
|
||||||
| GR-019 | Parity-blocked readiness load (gating without capability closure) | `logs/taskitem_runs/challenging_fullstack_multifile_20260226_r7/summary.json`, `logs/taskitem_runs/challenging_subset_prod_20260226_r7/summary.json`, `docs/sprint225_227_execution_tracker_2026-02-26.md` | `partial` | Sprint 225-227 reduced blocked parity load from `6` to `0` on both tracked hard catalogs while keeping unresolved divergence at `0`. This closes immediate safety debt for current corpora, but robustness is still contingent on pattern-driven repair classes. | Generalize repairs beyond queue-shaped transpile outputs |
|
| GR-019 | Parity-blocked readiness load (gating without capability closure) | `logs/taskitem_runs/challenging_fullstack_multifile_20260226_r7/summary.json`, `logs/taskitem_runs/challenging_subset_prod_20260226_r7/summary.json`, `docs/sprint225_227_execution_tracker_2026-02-26.md` | `partial` | Sprint 225-227 reduced blocked parity load from `6` to `0` on both tracked hard catalogs while keeping unresolved divergence at `0`. This closes immediate safety debt for current corpora, but robustness is still contingent on pattern-driven repair classes. | Generalize repairs beyond queue-shaped transpile outputs |
|
||||||
| GR-020 | Semantic fallback overuse masks weak native decomposition | `logs/taskitem_runs/TEST_ONLY_sprint236_semantic_fallback_audit_20260226/semantic_fallback_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint237_semantic_fallback_gate_fail_20260226/semantic_fallback_budget_gate.json`, `logs/taskitem_runs/TEST_ONLY_sprint238_native_gate_fail_20260226.json`, `logs/taskitem_runs/TEST_ONLY_sprint239_semantic_fallback_audit_20260226/semantic_fallback_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint240_retry_20260226_145631/00_summary.json`, `docs/sprint236_execution_tracker_2026-02-26.md`, `docs/sprint237_execution_tracker_2026-02-26.md`, `docs/sprint238_execution_tracker_2026-02-26.md`, `docs/sprint239_execution_tracker_2026-02-26.md`, `docs/sprint240_execution_tracker_2026-02-26.md` | `partial` | Sprint 236 added deterministic fallback-gap auditing with tool + constraint metadata. Sprint 237 added fallback-budget hard gate (`max-fallback-rate`) and produced expected fail/pass artifacts. Sprint 238 added native decomposition hard gate in pipeline (`min task count`, `min semantic signal count`) so weak native generation can be blocked before fallback masking. Sprint 239 added native reason enrichment and improved semantic signal density (`0 -> 6`). Sprint 240 added native decomposition retry with explicit minimum-task policy; on current sample retry attempted but did not improve depth (`2 -> 2`). | Native decomposition task-depth upgrade (increase native task granularity beyond 2) |
|
| GR-020 | Semantic fallback overuse masks weak native decomposition | `logs/taskitem_runs/TEST_ONLY_sprint236_semantic_fallback_audit_20260226/semantic_fallback_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint237_semantic_fallback_gate_fail_20260226/semantic_fallback_budget_gate.json`, `logs/taskitem_runs/TEST_ONLY_sprint238_native_gate_fail_20260226.json`, `logs/taskitem_runs/TEST_ONLY_sprint239_semantic_fallback_audit_20260226/semantic_fallback_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint240_retry_20260226_145631/00_summary.json`, `docs/sprint236_execution_tracker_2026-02-26.md`, `docs/sprint237_execution_tracker_2026-02-26.md`, `docs/sprint238_execution_tracker_2026-02-26.md`, `docs/sprint239_execution_tracker_2026-02-26.md`, `docs/sprint240_execution_tracker_2026-02-26.md` | `partial` | Sprint 236 added deterministic fallback-gap auditing with tool + constraint metadata. Sprint 237 added fallback-budget hard gate (`max-fallback-rate`) and produced expected fail/pass artifacts. Sprint 238 added native decomposition hard gate in pipeline (`min task count`, `min semantic signal count`) so weak native generation can be blocked before fallback masking. Sprint 239 added native reason enrichment and improved semantic signal density (`0 -> 6`). Sprint 240 added native decomposition retry with explicit minimum-task policy; on current sample retry attempted but did not improve depth (`2 -> 2`). | Native decomposition task-depth upgrade (increase native task granularity beyond 2) |
|
||||||
| GR-021 | Impact-specific native decomposition coverage not enforced uniformly | `docs/native_decomposition_impact_list_2026-02-26.md`, `tools/mcp/profiles/native_decomposition_impact_profiles.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_fullstack_impact_20260226_150416/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_native_impact_aggregate_20260226/native_impact_coverage_aggregate.json`, `logs/taskitem_runs/TEST_ONLY_sprint242_impact_remediation_loop_20260226/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint243_impact_remediation_tasks_loop_20260226_r2/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_20260226_151651/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_enforce_20260226_151709/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint245_intrinsic_20260226_151835/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_20260226_153033/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_enforce_20260226_153045/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_20260226_153230/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_enforce_20260226_153241/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint248_rawsearch_20260226_153903/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint249_rawguard_20260226_154028/02ae_raw_candidate_search.json`, `logs/taskitem_runs/TEST_ONLY_sprint250_closure_ladder_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint251_rawvariants_20260226_154355/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint252_closure_ladder_policy_skip2_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint253_closure_ladder_batch_20260226/closure_ladder_batch_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint254_closure_ladder_batch_20260226_r2/raw_gap_backlog.json`, `logs/taskitem_runs/TEST_ONLY_sprint255_rawharden_20260226_161924/00_summary.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162423/00_summary.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162633/02ae_raw_candidate_search.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162924/01e_raw_top_gap_requirements_report.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_163148/02af_raw_adaptive_retry_score.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_163556/02ae_raw_signal_targeted_variants.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_163855/01f_raw_intrinsic_prompt_pack_report.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_165430/01g_raw_template_control_pack_report.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_165746/02ae_candidate_0_projected_report.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_170014/06_native_impact_coverage.json` | `partial` | Sprint 241-262 established profile gating plus multiple intrinsic guidance paths; sprint 263 added deterministic structural projection before scoring and showed uplift in score telemetry; sprint 264 aligned score/gate parity by persisting selected raw candidate tasks into generation artifacts. On the hard sample with projector ON, both score and gate now pass consistently (`best_failing_profile_count=0`, `native_impact_coverage.failing_profile_count=0`, task_count `6`). Residual risk is breadth: closure is validated on this hard sample path, not yet proven across broader hard catalogs. | Batch-validate structural projector + parity path across challenging catalogs and add hard fail when score/gate diverge |
|
| GR-021 | Impact-specific native decomposition coverage not enforced uniformly | `docs/native_decomposition_impact_list_2026-02-26.md`, `tools/mcp/profiles/native_decomposition_impact_profiles.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_fullstack_impact_20260226_150416/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint241_native_impact_aggregate_20260226/native_impact_coverage_aggregate.json`, `logs/taskitem_runs/TEST_ONLY_sprint242_impact_remediation_loop_20260226/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint243_impact_remediation_tasks_loop_20260226_r2/remediation_loop_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_20260226_151651/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint244_autofill_enforce_20260226_151709/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint245_intrinsic_20260226_151835/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_20260226_153033/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint246_multishot_enforce_20260226_153045/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_20260226_153230/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint247_singleshot_enforce_20260226_153241/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint248_rawsearch_20260226_153903/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint249_rawguard_20260226_154028/02ae_raw_candidate_search.json`, `logs/taskitem_runs/TEST_ONLY_sprint250_closure_ladder_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint251_rawvariants_20260226_154355/00_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint252_closure_ladder_policy_skip2_20260226/closure_ladder_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint253_closure_ladder_batch_20260226/closure_ladder_batch_summary.json`, `logs/taskitem_runs/TEST_ONLY_sprint254_closure_ladder_batch_20260226_r2/raw_gap_backlog.json`, `logs/taskitem_runs/TEST_ONLY_sprint255_rawharden_20260226_161924/00_summary.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162423/00_summary.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162633/02ae_raw_candidate_search.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_162924/01e_raw_top_gap_requirements_report.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_163148/02af_raw_adaptive_retry_score.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_163556/02ae_raw_signal_targeted_variants.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_163855/01f_raw_intrinsic_prompt_pack_report.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_165430/01g_raw_template_control_pack_report.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_165746/02ae_candidate_0_projected_report.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_170014/06_native_impact_coverage.json`, `logs/taskitem_runs/01a_fallback_intake_spec_20260226_170301/00_summary.json` | `partial` | Sprint 241-262 established profile gating plus multiple intrinsic guidance paths; sprint 263 added deterministic structural projection before scoring; sprint 264 aligned score/gate parity by persisting selected raw candidate tasks into generation artifacts; sprint 265 added explicit `native_raw_score_gate_parity` packet and enforcement mode. On the hard sample with projector ON, score and gate both pass and parity packet reports `pass=true` under enforce mode. Residual risk is breadth: this closure is proven on the hard sample path but not yet validated across broader challenging catalogs. | Batch-validate projector+parity path across challenging catalogs and promote parity enforcement to default when stability is confirmed |
|
||||||
|
|
||||||
## What Was Covered Today (Sprints 175-184)
|
## What Was Covered Today (Sprints 175-184)
|
||||||
|
|
||||||
|
|||||||
@@ -556,3 +556,15 @@
|
|||||||
- `best_failing_profile_count=0`
|
- `best_failing_profile_count=0`
|
||||||
- `native_impact_coverage.failing_profile_count=0`
|
- `native_impact_coverage.failing_profile_count=0`
|
||||||
- scorer and gate are aligned.
|
- scorer and gate are aligned.
|
||||||
|
|
||||||
|
## Sprint 265 Added (Same Day)
|
||||||
|
|
||||||
|
- Added explicit score/gate parity telemetry + enforcement:
|
||||||
|
- `WSTONE_NATIVE_RAW_SCORE_GATE_PARITY_ENFORCE`
|
||||||
|
- summary packet:
|
||||||
|
- `native_raw_score_gate_parity`
|
||||||
|
- Dated validation artifact:
|
||||||
|
- `logs/taskitem_runs/01a_fallback_intake_spec_20260226_170301/00_summary.json`
|
||||||
|
- Current measured result on hard sample:
|
||||||
|
- parity packet reports `pass=true`
|
||||||
|
- enforcement enabled and run succeeds (`rc=0`).
|
||||||
|
|||||||
33
docs/sprint265_execution_tracker_2026-02-26.md
Normal file
33
docs/sprint265_execution_tracker_2026-02-26.md
Normal file
@@ -0,0 +1,33 @@
|
|||||||
|
# Sprint 265 Execution Tracker - 2026-02-26
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
- `sprint265_plan.md`
|
||||||
|
|
||||||
|
## Implemented
|
||||||
|
|
||||||
|
Pipeline parity packet/enforcement in `tools/mcp/run_sprint_taskitem_pipeline.sh`:
|
||||||
|
- Added:
|
||||||
|
- `WSTONE_NATIVE_RAW_SCORE_GATE_PARITY_ENFORCE`
|
||||||
|
- Added summary packet:
|
||||||
|
- `native_raw_score_gate_parity`
|
||||||
|
- `enabled`
|
||||||
|
- `pass`
|
||||||
|
- `raw_failing_profile_count`
|
||||||
|
- `gate_failing_profile_count`
|
||||||
|
- Enforcement behavior:
|
||||||
|
- when enabled, pipeline exits `19` if raw score and coverage gate diverge.
|
||||||
|
|
||||||
|
## Validation Artifacts
|
||||||
|
|
||||||
|
Parity-enforced run:
|
||||||
|
- `logs/taskitem_runs/01a_fallback_intake_spec_20260226_170301/00_summary.json`
|
||||||
|
|
||||||
|
Observed result on hard sample:
|
||||||
|
- `native_raw_candidate_search.best_failing_profile_count=0`
|
||||||
|
- `native_impact_coverage.failing_profile_count=0`
|
||||||
|
- `native_raw_score_gate_parity.pass=true`
|
||||||
|
- enforcement mode enabled and run succeeded (`rc=0`)
|
||||||
|
|
||||||
|
## Explicit Completion Signal
|
||||||
|
|
||||||
|
- Sprint 265: `DONE` (implemented + validated parity enforcement)
|
||||||
6
editor/src/Sprint265IntegrationSummary.h
Normal file
6
editor/src/Sprint265IntegrationSummary.h
Normal file
@@ -0,0 +1,6 @@
|
|||||||
|
#pragma once
|
||||||
|
|
||||||
|
// Sprint 265 integration summary:
|
||||||
|
// - Added native_raw_score_gate_parity telemetry packet.
|
||||||
|
// - Added optional hard-fail policy for score/gate divergence.
|
||||||
|
// - Parity enforcement now protects against silent drift between raw scorer and coverage gate.
|
||||||
10
sprint265_plan.md
Normal file
10
sprint265_plan.md
Normal file
@@ -0,0 +1,10 @@
|
|||||||
|
# Sprint 265 Plan: Raw Score/Gate Parity Packet and Enforcement
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
Make raw-score vs coverage-gate agreement a first-class runtime invariant with explicit telemetry and optional hard-fail enforcement.
|
||||||
|
|
||||||
|
## Steps
|
||||||
|
- Step 2372: Add native raw score/gate parity packet in pipeline summary.
|
||||||
|
- Step 2373: Add enforcement control that fails on divergence.
|
||||||
|
- Step 2374: Validate parity enforcement on hard sample with projector path.
|
||||||
|
- Step 2375: Add `Sprint265IntegrationSummary.h` and execution tracker.
|
||||||
@@ -42,6 +42,7 @@ NATIVE_DECOMP_TARGET_MIN_TASKS="${WSTONE_NATIVE_DECOMP_TARGET_MIN_TASKS:-5}"
|
|||||||
NATIVE_IMPACT_COVERAGE_GATE="${WSTONE_NATIVE_IMPACT_COVERAGE_GATE:-0}"
|
NATIVE_IMPACT_COVERAGE_GATE="${WSTONE_NATIVE_IMPACT_COVERAGE_GATE:-0}"
|
||||||
NATIVE_IMPACT_COVERAGE_ENFORCE="${WSTONE_NATIVE_IMPACT_COVERAGE_ENFORCE:-0}"
|
NATIVE_IMPACT_COVERAGE_ENFORCE="${WSTONE_NATIVE_IMPACT_COVERAGE_ENFORCE:-0}"
|
||||||
NATIVE_IMPACT_COVERAGE_PROFILES="${WSTONE_NATIVE_IMPACT_COVERAGE_PROFILES:-$ROOT_DIR/tools/mcp/profiles/native_decomposition_impact_profiles.json}"
|
NATIVE_IMPACT_COVERAGE_PROFILES="${WSTONE_NATIVE_IMPACT_COVERAGE_PROFILES:-$ROOT_DIR/tools/mcp/profiles/native_decomposition_impact_profiles.json}"
|
||||||
|
NATIVE_RAW_SCORE_GATE_PARITY_ENFORCE="${WSTONE_NATIVE_RAW_SCORE_GATE_PARITY_ENFORCE:-0}"
|
||||||
NATIVE_PROFILE_AUTOFILL="${WSTONE_NATIVE_PROFILE_AUTOFILL:-0}"
|
NATIVE_PROFILE_AUTOFILL="${WSTONE_NATIVE_PROFILE_AUTOFILL:-0}"
|
||||||
NATIVE_PROFILE_AUTOFILL_MAX_TASKS="${WSTONE_NATIVE_PROFILE_AUTOFILL_MAX_TASKS:-12}"
|
NATIVE_PROFILE_AUTOFILL_MAX_TASKS="${WSTONE_NATIVE_PROFILE_AUTOFILL_MAX_TASKS:-12}"
|
||||||
NATIVE_INTRINSIC_BOOST="${WSTONE_NATIVE_INTRINSIC_BOOST:-0}"
|
NATIVE_INTRINSIC_BOOST="${WSTONE_NATIVE_INTRINSIC_BOOST:-0}"
|
||||||
@@ -124,6 +125,7 @@ NATIVE_DECOMP_GATE_JSON='{}'
|
|||||||
NATIVE_REASON_ENRICHMENT_JSON='{}'
|
NATIVE_REASON_ENRICHMENT_JSON='{}'
|
||||||
NATIVE_DECOMP_RETRY_JSON='{}'
|
NATIVE_DECOMP_RETRY_JSON='{}'
|
||||||
NATIVE_IMPACT_COVERAGE_JSON='{}'
|
NATIVE_IMPACT_COVERAGE_JSON='{}'
|
||||||
|
NATIVE_RAW_SCORE_GATE_PARITY_JSON='{}'
|
||||||
NATIVE_PROFILE_AUTOFILL_JSON='{}'
|
NATIVE_PROFILE_AUTOFILL_JSON='{}'
|
||||||
NATIVE_INTRINSIC_BOOST_JSON='{}'
|
NATIVE_INTRINSIC_BOOST_JSON='{}'
|
||||||
NATIVE_SINGLESHOT_PROFILE_SHAPE_JSON='{}'
|
NATIVE_SINGLESHOT_PROFILE_SHAPE_JSON='{}'
|
||||||
@@ -1160,6 +1162,31 @@ if [[ "$NATIVE_IMPACT_COVERAGE_GATE" == "1" ]]; then
|
|||||||
SUMMARY_JSON="$(printf '%s' "$SUMMARY_JSON" | jq --argjson nic "$NATIVE_IMPACT_COVERAGE_JSON" '.native_impact_coverage = $nic')"
|
SUMMARY_JSON="$(printf '%s' "$SUMMARY_JSON" | jq --argjson nic "$NATIVE_IMPACT_COVERAGE_JSON" '.native_impact_coverage = $nic')"
|
||||||
fi
|
fi
|
||||||
|
|
||||||
|
if [[ "$NATIVE_RAW_CANDIDATE_SEARCH" == "1" && "$NATIVE_IMPACT_COVERAGE_GATE" == "1" ]]; then
|
||||||
|
raw_fail_count="$(printf '%s' "$NATIVE_RAW_CANDIDATE_SEARCH_JSON" | jq '.best_failing_profile_count // null')"
|
||||||
|
gate_fail_count="$(printf '%s' "$NATIVE_IMPACT_COVERAGE_JSON" | jq '.failing_profile_count // null')"
|
||||||
|
parity_pass=false
|
||||||
|
if [[ "$raw_fail_count" == "$gate_fail_count" ]]; then
|
||||||
|
parity_pass=true
|
||||||
|
fi
|
||||||
|
NATIVE_RAW_SCORE_GATE_PARITY_JSON="$(jq -nc \
|
||||||
|
--argjson enabled true \
|
||||||
|
--argjson pass "$parity_pass" \
|
||||||
|
--argjson raw_failing_profile_count "$raw_fail_count" \
|
||||||
|
--argjson gate_failing_profile_count "$gate_fail_count" \
|
||||||
|
'{enabled:$enabled, pass:$pass, raw_failing_profile_count:$raw_failing_profile_count, gate_failing_profile_count:$gate_failing_profile_count}')"
|
||||||
|
SUMMARY_JSON="$(printf '%s' "$SUMMARY_JSON" | jq --argjson p "$NATIVE_RAW_SCORE_GATE_PARITY_JSON" '.native_raw_score_gate_parity = $p')"
|
||||||
|
if [[ "$NATIVE_RAW_SCORE_GATE_PARITY_ENFORCE" == "1" && "$parity_pass" != "true" ]]; then
|
||||||
|
printf '%s\n' "$SUMMARY_JSON" > "$OUT_DIR/00_summary.json"
|
||||||
|
echo "error: raw score/gate parity check failed" >&2
|
||||||
|
echo "error: raw_failing_profile_count=$raw_fail_count gate_failing_profile_count=$gate_fail_count" >&2
|
||||||
|
exit 19
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
NATIVE_RAW_SCORE_GATE_PARITY_JSON='{"enabled":false}'
|
||||||
|
SUMMARY_JSON="$(printf '%s' "$SUMMARY_JSON" | jq --argjson p "$NATIVE_RAW_SCORE_GATE_PARITY_JSON" '.native_raw_score_gate_parity = $p')"
|
||||||
|
fi
|
||||||
|
|
||||||
if [[ "$CALIBRATE_AFTER_RUN" == "1" ]]; then
|
if [[ "$CALIBRATE_AFTER_RUN" == "1" ]]; then
|
||||||
CALIBRATION_OUT_DIR="$OUT_DIR/calibration"
|
CALIBRATION_OUT_DIR="$OUT_DIR/calibration"
|
||||||
CALIBRATION_JSON="$(python3 "$ROOT_DIR/tools/mcp/analyze_taskitem_calibration.py" \
|
CALIBRATION_JSON="$(python3 "$ROOT_DIR/tools/mcp/analyze_taskitem_calibration.py" \
|
||||||
|
|||||||
Reference in New Issue
Block a user