5.9 KiB
Evaluation Policy
Purpose
This document defines how bounded decision gates should be evaluated inside
whetstone_RSA.
The central rule is:
Do not optimize for raw accuracy in isolation.
Optimize for final correctness at acceptable latency and cost, while keeping silent wrong decisions below the gate's risk budget.
Why Raw Accuracy Is Insufficient
A gate with 70% raw top-1 accuracy may still be useful if:
- it handles easy cases confidently
- it abstains on risky cases
- retries recover a meaningful share of the abstentions
- downstream deterministic checks catch many residual mistakes
A gate with 97% raw accuracy may still be a poor deployment choice if:
- it requires 10x to 16x model size
- latency becomes unacceptable
- it blocks parallel deployment
- the extra accuracy only reduces low-cost mistakes
For bounded systems, the real question is:
What is the expected cost of using this gate in the pipeline?
Required Metrics
Every gate evaluation should report these metrics:
raw_accuracyaccept_rateaccuracy_on_acceptedabstain_rateretry_recovery_rateescalation_ratesilent_error_rateavg_latency_msavg_compute_cost
Optional but recommended:
- calibration error
- confusion concentration
- class-weighted error
- deterministic-check catch rate
- rollback or repair rate
Definitions
Raw Accuracy
Top-1 accuracy on the full evaluation set with no abstention policy.
Useful for research. Not enough for deployment.
Accept Rate
Fraction of inputs the gate handles automatically without abstaining.
Accuracy On Accepted
Accuracy measured only on predictions that pass the acceptance threshold.
This is often more important than raw accuracy.
Retry Recovery Rate
Fraction of abstained or low-confidence cases that become correct after:
- adding richer context
- running a deterministic pre-pass
- using a larger specialist
Silent Error Rate
Wrong decisions that are accepted automatically and are not caught by later deterministic checks before they can do damage.
This is usually the key deployment risk metric.
Gate Risk Tiers
Every gate should be assigned a risk tier before model sizing decisions are made.
Tier 1: Low Risk
Characteristics:
- mistakes are easy to reverse
- downstream deterministic checks exist
- retries are cheap
- wrong branch selection does not poison later stages
Policy:
- lower raw accuracy can be acceptable
- moderate abstain rate is fine
- optimize for cheap throughput
Tier 2: Medium Risk
Characteristics:
- mistakes cause wasted work or moderate workflow churn
- some downstream correction exists, but not always
- retries have non-trivial cost
Policy:
- accepted predictions need stronger precision
- confidence thresholds should be conservative
- escalation should be available
Tier 3: High Risk
Characteristics:
- mistakes silently corrupt planning or structure
- correction is costly or delayed
- wrong decisions can contaminate later stages
Policy:
- silent error rate must be very low
- abstain aggressively
- escalate early
- raw accuracy matters less than safe acceptance behavior
Decision Rule
For a given gate, accept a smaller model if it wins on the system frontier:
- lower average cost
- lower or equal silent error rate
- acceptable final correctness after retry and escalation
- acceptable latency
Do not move to a larger model just because it raises raw accuracy.
Move to a larger model only if:
- silent errors remain too high after thresholding
- retries do not recover enough cases
- the gate is structurally valid and still under-capacity
Pre-Scale Checks
Before scaling a gate from tiny to medium or large, check:
- is the output schema well-defined?
- are the labels under-enumerated?
- are there ambiguous labels that should be split?
- is important context missing from the input encoding?
- can the gate be decomposed into two simpler gates?
- should part of the logic be deterministic instead?
Use docs/gate_failure_taxonomy.md to classify which of these checks should be
prioritized for a given weak gate.
If any of those fail, bigger models may just hide a design problem.
Recommended Evaluation Loop
- Train and measure raw accuracy.
- Calibrate confidence thresholds.
- Measure
accuracy_on_acceptedandsilent_error_rate. - Add deterministic checks and retry path.
- Measure final correctness and average cost.
- Only then compare larger model tiers.
Working Heuristic
Use this default posture until domain-specific policy replaces it:
- low-risk gates: a base model in the
70-85%range may still be deployable with abstain and retry - medium-risk gates: target stronger accepted precision, not just higher raw accuracy
- high-risk gates: prioritize low silent-error rate over broad auto-accept coverage
These are not hard thresholds. They are a reminder that the deployment objective is system utility, not benchmark vanity.
WhetstoneDSL Case Study
Current tiny-tier WhetstoneDSL results already show why this policy is needed:
confidence_tierreaches100%, but is essentially a formula taskverification_typereaches roughly85%worker_typereaches roughly88.5%prereq_opis flat at roughly69.4%automatabilityis flat at roughly69.6%
The correct interpretation is not:
- "
69%is bad, so scale immediately"
The correct interpretation is:
- determine whether those
69%gates can be made safe with acceptance thresholds, decomposition, retries, and deterministic guards - then decide whether larger specialists are economically justified
What To Log
Each live decision should eventually log:
- gate id
- model tier
- confidence
- accepted or abstained
- retry used or not
- escalation target
- final chosen action
- deterministic validation outcome
- final human or system override outcome
That turns deployment into a continuous training and policy dataset instead of a one-time benchmark.