Files
whetstone_RSA/docs/evaluation_policy.md
2026-03-31 22:50:40 -06:00

5.9 KiB

Evaluation Policy

Purpose

This document defines how bounded decision gates should be evaluated inside whetstone_RSA.

The central rule is:

Do not optimize for raw accuracy in isolation.

Optimize for final correctness at acceptable latency and cost, while keeping silent wrong decisions below the gate's risk budget.

Why Raw Accuracy Is Insufficient

A gate with 70% raw top-1 accuracy may still be useful if:

  • it handles easy cases confidently
  • it abstains on risky cases
  • retries recover a meaningful share of the abstentions
  • downstream deterministic checks catch many residual mistakes

A gate with 97% raw accuracy may still be a poor deployment choice if:

  • it requires 10x to 16x model size
  • latency becomes unacceptable
  • it blocks parallel deployment
  • the extra accuracy only reduces low-cost mistakes

For bounded systems, the real question is:

What is the expected cost of using this gate in the pipeline?

Required Metrics

Every gate evaluation should report these metrics:

  • raw_accuracy
  • accept_rate
  • accuracy_on_accepted
  • abstain_rate
  • retry_recovery_rate
  • escalation_rate
  • silent_error_rate
  • avg_latency_ms
  • avg_compute_cost

Optional but recommended:

  • calibration error
  • confusion concentration
  • class-weighted error
  • deterministic-check catch rate
  • rollback or repair rate

Definitions

Raw Accuracy

Top-1 accuracy on the full evaluation set with no abstention policy.

Useful for research. Not enough for deployment.

Accept Rate

Fraction of inputs the gate handles automatically without abstaining.

Accuracy On Accepted

Accuracy measured only on predictions that pass the acceptance threshold.

This is often more important than raw accuracy.

Retry Recovery Rate

Fraction of abstained or low-confidence cases that become correct after:

  • adding richer context
  • running a deterministic pre-pass
  • using a larger specialist

Silent Error Rate

Wrong decisions that are accepted automatically and are not caught by later deterministic checks before they can do damage.

This is usually the key deployment risk metric.

Gate Risk Tiers

Every gate should be assigned a risk tier before model sizing decisions are made.

Tier 1: Low Risk

Characteristics:

  • mistakes are easy to reverse
  • downstream deterministic checks exist
  • retries are cheap
  • wrong branch selection does not poison later stages

Policy:

  • lower raw accuracy can be acceptable
  • moderate abstain rate is fine
  • optimize for cheap throughput

Tier 2: Medium Risk

Characteristics:

  • mistakes cause wasted work or moderate workflow churn
  • some downstream correction exists, but not always
  • retries have non-trivial cost

Policy:

  • accepted predictions need stronger precision
  • confidence thresholds should be conservative
  • escalation should be available

Tier 3: High Risk

Characteristics:

  • mistakes silently corrupt planning or structure
  • correction is costly or delayed
  • wrong decisions can contaminate later stages

Policy:

  • silent error rate must be very low
  • abstain aggressively
  • escalate early
  • raw accuracy matters less than safe acceptance behavior

Decision Rule

For a given gate, accept a smaller model if it wins on the system frontier:

  • lower average cost
  • lower or equal silent error rate
  • acceptable final correctness after retry and escalation
  • acceptable latency

Do not move to a larger model just because it raises raw accuracy.

Move to a larger model only if:

  • silent errors remain too high after thresholding
  • retries do not recover enough cases
  • the gate is structurally valid and still under-capacity

Pre-Scale Checks

Before scaling a gate from tiny to medium or large, check:

  • is the output schema well-defined?
  • are the labels under-enumerated?
  • are there ambiguous labels that should be split?
  • is important context missing from the input encoding?
  • can the gate be decomposed into two simpler gates?
  • should part of the logic be deterministic instead?

Use docs/gate_failure_taxonomy.md to classify which of these checks should be prioritized for a given weak gate.

If any of those fail, bigger models may just hide a design problem.

  1. Train and measure raw accuracy.
  2. Calibrate confidence thresholds.
  3. Measure accuracy_on_accepted and silent_error_rate.
  4. Add deterministic checks and retry path.
  5. Measure final correctness and average cost.
  6. Only then compare larger model tiers.

Working Heuristic

Use this default posture until domain-specific policy replaces it:

  • low-risk gates: a base model in the 70-85% range may still be deployable with abstain and retry
  • medium-risk gates: target stronger accepted precision, not just higher raw accuracy
  • high-risk gates: prioritize low silent-error rate over broad auto-accept coverage

These are not hard thresholds. They are a reminder that the deployment objective is system utility, not benchmark vanity.

WhetstoneDSL Case Study

Current tiny-tier WhetstoneDSL results already show why this policy is needed:

  • confidence_tier reaches 100%, but is essentially a formula task
  • verification_type reaches roughly 85%
  • worker_type reaches roughly 88.5%
  • prereq_op is flat at roughly 69.4%
  • automatability is flat at roughly 69.6%

The correct interpretation is not:

  • "69% is bad, so scale immediately"

The correct interpretation is:

  • determine whether those 69% gates can be made safe with acceptance thresholds, decomposition, retries, and deterministic guards
  • then decide whether larger specialists are economically justified

What To Log

Each live decision should eventually log:

  • gate id
  • model tier
  • confidence
  • accepted or abstained
  • retry used or not
  • escalation target
  • final chosen action
  • deterministic validation outcome
  • final human or system override outcome

That turns deployment into a continuous training and policy dataset instead of a one-time benchmark.