Initial whetstone_RSA architecture and C++ scaffold
This commit is contained in:
171
docs/gate_rubric.md
Normal file
171
docs/gate_rubric.md
Normal file
@@ -0,0 +1,171 @@
|
||||
# Gate Rubric
|
||||
|
||||
## Purpose
|
||||
|
||||
This document gives a provisional deployment rubric for the current WhetstoneDSL
|
||||
specialist gates.
|
||||
|
||||
The goal is not to rank models by benchmark score. The goal is to decide, per gate:
|
||||
|
||||
- deploy now
|
||||
- deploy with guardrails
|
||||
- wait for larger-tier comparison
|
||||
- redesign the gate before scaling
|
||||
|
||||
## Rubric Buckets
|
||||
|
||||
### Deploy Now
|
||||
|
||||
Use when:
|
||||
|
||||
- the gate is effectively deterministic or near-deterministic
|
||||
- wrong decisions are easy to catch or reverse
|
||||
- the current model already clears the practical quality bar
|
||||
|
||||
### Deploy With Guardrails
|
||||
|
||||
Use when:
|
||||
|
||||
- the gate is useful at the current size
|
||||
- raw accuracy is not enough for blind auto-accept
|
||||
- confidence thresholding, retry, or deterministic checks can contain risk
|
||||
|
||||
### Wait For Large Run
|
||||
|
||||
Use when:
|
||||
|
||||
- the current tiny-tier result is too weak for safe deployment
|
||||
- the gate definition appears defensible
|
||||
- larger-tier quality data is actively being generated
|
||||
|
||||
### Redesign Before Scaling
|
||||
|
||||
Use when:
|
||||
|
||||
- the label space appears structurally muddy
|
||||
- the gate may really be multiple smaller gates
|
||||
- important input context is missing
|
||||
- bigger models would mostly compensate for a bad formulation
|
||||
|
||||
## Current WhetstoneDSL Gates
|
||||
|
||||
### `confidence_tier`
|
||||
|
||||
- Tiny-tier result: `100%`
|
||||
- Provisional bucket: `Deploy Now`
|
||||
- Rationale:
|
||||
- this is very close to a formula task
|
||||
- it should be treated as a deterministic baseline whenever possible
|
||||
- specialist use is acceptable, but only as a convenience layer over a known rule
|
||||
|
||||
### `verification_type`
|
||||
|
||||
- Tiny-tier result: about `84.9%`
|
||||
- Provisional bucket: `Deploy With Guardrails`
|
||||
- Rationale:
|
||||
- bounded 5-way task with real signal at the tiny tier
|
||||
- likely useful if accepted predictions have high precision
|
||||
- reversible in many workflows because verification type can be validated or
|
||||
corrected downstream
|
||||
- Recommended next step:
|
||||
- evaluate confidence-thresholded precision
|
||||
- inspect confusion concentration, especially adjacent classes
|
||||
|
||||
### `worker_type`
|
||||
|
||||
- Tiny-tier result: about `88.5%`
|
||||
- Provisional bucket: `Deploy With Guardrails`
|
||||
- Rationale:
|
||||
- strong enough to be operationally useful if low-confidence cases abstain
|
||||
- wrong worker selection is often recoverable
|
||||
- likely a good candidate for first-pass routing plus retry
|
||||
- Recommended next step:
|
||||
- measure accepted precision and abstain coverage
|
||||
- add a cheap retry path with richer context
|
||||
|
||||
### `prereq_op`
|
||||
|
||||
- Tiny-tier result: about `69.4%`, flat across steps 1000-4000
|
||||
- Provisional bucket: `Redesign Before Scaling`
|
||||
- Rationale:
|
||||
- the current 4-way label is really a composition of two latent yes/no decisions:
|
||||
`needs architect review?` and `needs manual approval?`
|
||||
- the current generator maps those two axes into one of:
|
||||
`standard`, `needs_review`, `needs_approval`, `full_gates`
|
||||
- that means the model is being asked to infer a product label from short text,
|
||||
even though the natural decision structure is factorized
|
||||
- labels are heuristic and keyword-driven, which likely adds noise
|
||||
- scaling the model may help, but it may also just memorize fuzzy boundaries
|
||||
- Why it is hard:
|
||||
- short descriptions often under-specify whether a task is architectural,
|
||||
production-sensitive, both, or neither
|
||||
- many cues are policy-level rather than lexical
|
||||
- some examples mix "phase/integration/framework" signals with release/security
|
||||
signals, making the combined class boundary weak
|
||||
- the downstream action is itself compositional, which suggests the predictor
|
||||
should also be compositional
|
||||
- Recommended decomposition:
|
||||
- gate A: `requires_architect_review` → `yes/no`
|
||||
- gate B: `requires_manual_approval` → `yes/no`
|
||||
- deterministic combiner:
|
||||
- `00` -> `standard`
|
||||
- `10` -> `needs_review`
|
||||
- `01` -> `needs_approval`
|
||||
- `11` -> `full_gates`
|
||||
- Why decomposition is attractive:
|
||||
- each model solves a cleaner binary problem
|
||||
- outputs align directly with the policy semantics
|
||||
- confidence can be calibrated independently per axis
|
||||
- one axis may be deterministic or near-deterministic in some contexts
|
||||
- What to test before scaling:
|
||||
- binary performance for each axis
|
||||
- whether richer structured inputs outperform plain step text
|
||||
- whether some approval decisions should be deterministic policy checks instead
|
||||
|
||||
### `automatability`
|
||||
|
||||
- Tiny-tier result: about `69.6%`, flat across steps 1000-4000
|
||||
- Provisional bucket: `Wait For Large Run` and `Redesign Before Scaling`
|
||||
- Rationale:
|
||||
- this gate is conceptually important, but the label boundaries are likely moving
|
||||
targets because they depend on what tools and templates the system already has
|
||||
- adjacent classes are semantically close
|
||||
- the cost of some confusions is very asymmetric
|
||||
- Why it is hard:
|
||||
- the classes are not purely descriptive; they are partly roadmap-dependent
|
||||
- a node can move from `llm` to `specialist` or `template` as tooling improves
|
||||
- the difference between `template`, `specialist`, `slm`, and `llm` is partly an
|
||||
operational routing policy, not just an intrinsic property of the node
|
||||
- Recommended next step:
|
||||
- wait for current large runs to finish
|
||||
- inspect confusion matrix, not just headline accuracy
|
||||
- consider a hierarchical formulation
|
||||
- Possible decomposition:
|
||||
- gate A: `is deterministic/template sufficient?`
|
||||
- gate B: `if learned, is output bounded or open-vocabulary?`
|
||||
- gate C: `if open-vocabulary, does small-model routing suffice or is full LLM needed?`
|
||||
|
||||
## Immediate Operational Policy
|
||||
|
||||
Until the larger runs complete, use this posture:
|
||||
|
||||
- `confidence_tier`: deterministic or auto-deploy
|
||||
- `verification_type`: candidate for guarded deployment
|
||||
- `worker_type`: candidate for guarded deployment
|
||||
- `prereq_op`: do not blindly scale; test decomposed formulation first
|
||||
- `automatability`: keep training and collect richer evidence before committing
|
||||
|
||||
## Main Lesson
|
||||
|
||||
When a bounded gate performs poorly, the first question should not be:
|
||||
|
||||
"How much bigger should the model be?"
|
||||
|
||||
The first questions should be:
|
||||
|
||||
- is the gate actually one gate?
|
||||
- is the label space stable?
|
||||
- is the input missing policy-critical structure?
|
||||
- should part of the logic be deterministic?
|
||||
|
||||
That is the point of this rubric.
|
||||
Reference in New Issue
Block a user