Files
whetstone_RSA/docs/evaluation_policy.md
2026-03-31 22:50:40 -06:00

228 lines
5.9 KiB
Markdown

# Evaluation Policy
## Purpose
This document defines how bounded decision gates should be evaluated inside
`whetstone_RSA`.
The central rule is:
Do not optimize for raw accuracy in isolation.
Optimize for final correctness at acceptable latency and cost, while keeping
silent wrong decisions below the gate's risk budget.
## Why Raw Accuracy Is Insufficient
A gate with `70%` raw top-1 accuracy may still be useful if:
- it handles easy cases confidently
- it abstains on risky cases
- retries recover a meaningful share of the abstentions
- downstream deterministic checks catch many residual mistakes
A gate with `97%` raw accuracy may still be a poor deployment choice if:
- it requires 10x to 16x model size
- latency becomes unacceptable
- it blocks parallel deployment
- the extra accuracy only reduces low-cost mistakes
For bounded systems, the real question is:
What is the expected cost of using this gate in the pipeline?
## Required Metrics
Every gate evaluation should report these metrics:
- `raw_accuracy`
- `accept_rate`
- `accuracy_on_accepted`
- `abstain_rate`
- `retry_recovery_rate`
- `escalation_rate`
- `silent_error_rate`
- `avg_latency_ms`
- `avg_compute_cost`
Optional but recommended:
- calibration error
- confusion concentration
- class-weighted error
- deterministic-check catch rate
- rollback or repair rate
## Definitions
### Raw Accuracy
Top-1 accuracy on the full evaluation set with no abstention policy.
Useful for research. Not enough for deployment.
### Accept Rate
Fraction of inputs the gate handles automatically without abstaining.
### Accuracy On Accepted
Accuracy measured only on predictions that pass the acceptance threshold.
This is often more important than raw accuracy.
### Retry Recovery Rate
Fraction of abstained or low-confidence cases that become correct after:
- adding richer context
- running a deterministic pre-pass
- using a larger specialist
### Silent Error Rate
Wrong decisions that are accepted automatically and are not caught by later
deterministic checks before they can do damage.
This is usually the key deployment risk metric.
## Gate Risk Tiers
Every gate should be assigned a risk tier before model sizing decisions are made.
### Tier 1: Low Risk
Characteristics:
- mistakes are easy to reverse
- downstream deterministic checks exist
- retries are cheap
- wrong branch selection does not poison later stages
Policy:
- lower raw accuracy can be acceptable
- moderate abstain rate is fine
- optimize for cheap throughput
### Tier 2: Medium Risk
Characteristics:
- mistakes cause wasted work or moderate workflow churn
- some downstream correction exists, but not always
- retries have non-trivial cost
Policy:
- accepted predictions need stronger precision
- confidence thresholds should be conservative
- escalation should be available
### Tier 3: High Risk
Characteristics:
- mistakes silently corrupt planning or structure
- correction is costly or delayed
- wrong decisions can contaminate later stages
Policy:
- silent error rate must be very low
- abstain aggressively
- escalate early
- raw accuracy matters less than safe acceptance behavior
## Decision Rule
For a given gate, accept a smaller model if it wins on the system frontier:
- lower average cost
- lower or equal silent error rate
- acceptable final correctness after retry and escalation
- acceptable latency
Do not move to a larger model just because it raises raw accuracy.
Move to a larger model only if:
- silent errors remain too high after thresholding
- retries do not recover enough cases
- the gate is structurally valid and still under-capacity
## Pre-Scale Checks
Before scaling a gate from tiny to medium or large, check:
- is the output schema well-defined?
- are the labels under-enumerated?
- are there ambiguous labels that should be split?
- is important context missing from the input encoding?
- can the gate be decomposed into two simpler gates?
- should part of the logic be deterministic instead?
Use `docs/gate_failure_taxonomy.md` to classify which of these checks should be
prioritized for a given weak gate.
If any of those fail, bigger models may just hide a design problem.
## Recommended Evaluation Loop
1. Train and measure raw accuracy.
2. Calibrate confidence thresholds.
3. Measure `accuracy_on_accepted` and `silent_error_rate`.
4. Add deterministic checks and retry path.
5. Measure final correctness and average cost.
6. Only then compare larger model tiers.
## Working Heuristic
Use this default posture until domain-specific policy replaces it:
- low-risk gates: a base model in the `70-85%` range may still be deployable with
abstain and retry
- medium-risk gates: target stronger accepted precision, not just higher raw accuracy
- high-risk gates: prioritize low silent-error rate over broad auto-accept coverage
These are not hard thresholds. They are a reminder that the deployment objective is
system utility, not benchmark vanity.
## WhetstoneDSL Case Study
Current tiny-tier WhetstoneDSL results already show why this policy is needed:
- `confidence_tier` reaches `100%`, but is essentially a formula task
- `verification_type` reaches roughly `85%`
- `worker_type` reaches roughly `88.5%`
- `prereq_op` is flat at roughly `69.4%`
- `automatability` is flat at roughly `69.6%`
The correct interpretation is not:
- "`69%` is bad, so scale immediately"
The correct interpretation is:
- determine whether those `69%` gates can be made safe with acceptance thresholds,
decomposition, retries, and deterministic guards
- then decide whether larger specialists are economically justified
## What To Log
Each live decision should eventually log:
- gate id
- model tier
- confidence
- accepted or abstained
- retry used or not
- escalation target
- final chosen action
- deterministic validation outcome
- final human or system override outcome
That turns deployment into a continuous training and policy dataset instead of a
one-time benchmark.