Files
whetstone_RSA/docs/gate_rubric.md

172 lines
6.1 KiB
Markdown
Raw Normal View History

# Gate Rubric
## Purpose
This document gives a provisional deployment rubric for the current WhetstoneDSL
specialist gates.
The goal is not to rank models by benchmark score. The goal is to decide, per gate:
- deploy now
- deploy with guardrails
- wait for larger-tier comparison
- redesign the gate before scaling
## Rubric Buckets
### Deploy Now
Use when:
- the gate is effectively deterministic or near-deterministic
- wrong decisions are easy to catch or reverse
- the current model already clears the practical quality bar
### Deploy With Guardrails
Use when:
- the gate is useful at the current size
- raw accuracy is not enough for blind auto-accept
- confidence thresholding, retry, or deterministic checks can contain risk
### Wait For Large Run
Use when:
- the current tiny-tier result is too weak for safe deployment
- the gate definition appears defensible
- larger-tier quality data is actively being generated
### Redesign Before Scaling
Use when:
- the label space appears structurally muddy
- the gate may really be multiple smaller gates
- important input context is missing
- bigger models would mostly compensate for a bad formulation
## Current WhetstoneDSL Gates
### `confidence_tier`
- Tiny-tier result: `100%`
- Provisional bucket: `Deploy Now`
- Rationale:
- this is very close to a formula task
- it should be treated as a deterministic baseline whenever possible
- specialist use is acceptable, but only as a convenience layer over a known rule
### `verification_type`
- Tiny-tier result: about `84.9%`
- Provisional bucket: `Deploy With Guardrails`
- Rationale:
- bounded 5-way task with real signal at the tiny tier
- likely useful if accepted predictions have high precision
- reversible in many workflows because verification type can be validated or
corrected downstream
- Recommended next step:
- evaluate confidence-thresholded precision
- inspect confusion concentration, especially adjacent classes
### `worker_type`
- Tiny-tier result: about `88.5%`
- Provisional bucket: `Deploy With Guardrails`
- Rationale:
- strong enough to be operationally useful if low-confidence cases abstain
- wrong worker selection is often recoverable
- likely a good candidate for first-pass routing plus retry
- Recommended next step:
- measure accepted precision and abstain coverage
- add a cheap retry path with richer context
### `prereq_op`
- Tiny-tier result: about `69.4%`, flat across steps 1000-4000
- Provisional bucket: `Redesign Before Scaling`
- Rationale:
- the current 4-way label is really a composition of two latent yes/no decisions:
`needs architect review?` and `needs manual approval?`
- the current generator maps those two axes into one of:
`standard`, `needs_review`, `needs_approval`, `full_gates`
- that means the model is being asked to infer a product label from short text,
even though the natural decision structure is factorized
- labels are heuristic and keyword-driven, which likely adds noise
- scaling the model may help, but it may also just memorize fuzzy boundaries
- Why it is hard:
- short descriptions often under-specify whether a task is architectural,
production-sensitive, both, or neither
- many cues are policy-level rather than lexical
- some examples mix "phase/integration/framework" signals with release/security
signals, making the combined class boundary weak
- the downstream action is itself compositional, which suggests the predictor
should also be compositional
- Recommended decomposition:
- gate A: `requires_architect_review``yes/no`
- gate B: `requires_manual_approval``yes/no`
- deterministic combiner:
- `00` -> `standard`
- `10` -> `needs_review`
- `01` -> `needs_approval`
- `11` -> `full_gates`
- Why decomposition is attractive:
- each model solves a cleaner binary problem
- outputs align directly with the policy semantics
- confidence can be calibrated independently per axis
- one axis may be deterministic or near-deterministic in some contexts
- What to test before scaling:
- binary performance for each axis
- whether richer structured inputs outperform plain step text
- whether some approval decisions should be deterministic policy checks instead
### `automatability`
- Tiny-tier result: about `69.6%`, flat across steps 1000-4000
- Provisional bucket: `Wait For Large Run` and `Redesign Before Scaling`
- Rationale:
- this gate is conceptually important, but the label boundaries are likely moving
targets because they depend on what tools and templates the system already has
- adjacent classes are semantically close
- the cost of some confusions is very asymmetric
- Why it is hard:
- the classes are not purely descriptive; they are partly roadmap-dependent
- a node can move from `llm` to `specialist` or `template` as tooling improves
- the difference between `template`, `specialist`, `slm`, and `llm` is partly an
operational routing policy, not just an intrinsic property of the node
- Recommended next step:
- wait for current large runs to finish
- inspect confusion matrix, not just headline accuracy
- consider a hierarchical formulation
- Possible decomposition:
- gate A: `is deterministic/template sufficient?`
- gate B: `if learned, is output bounded or open-vocabulary?`
- gate C: `if open-vocabulary, does small-model routing suffice or is full LLM needed?`
## Immediate Operational Policy
Until the larger runs complete, use this posture:
- `confidence_tier`: deterministic or auto-deploy
- `verification_type`: candidate for guarded deployment
- `worker_type`: candidate for guarded deployment
- `prereq_op`: do not blindly scale; test decomposed formulation first
- `automatability`: keep training and collect richer evidence before committing
## Main Lesson
When a bounded gate performs poorly, the first question should not be:
"How much bigger should the model be?"
The first questions should be:
- is the gate actually one gate?
- is the label space stable?
- is the input missing policy-critical structure?
- should part of the logic be deterministic?
That is the point of this rubric.