392 lines
10 KiB
Markdown
392 lines
10 KiB
Markdown
|
|
# Gate Failure Taxonomy
|
||
|
|
|
||
|
|
## Purpose
|
||
|
|
|
||
|
|
This document defines a first-pass taxonomy for why bounded decision gates fail and
|
||
|
|
how `whetstone_RSA` should reason about those failures.
|
||
|
|
|
||
|
|
The goal is to make gate diagnosis a library capability rather than an ad hoc human
|
||
|
|
judgment call.
|
||
|
|
|
||
|
|
This taxonomy is motivated by the `whetstone_DSL` case study, but it is written as
|
||
|
|
an RSA-library abstraction.
|
||
|
|
|
||
|
|
## Core Claim
|
||
|
|
|
||
|
|
When a bounded gate performs poorly, the default answer should not be "use a larger
|
||
|
|
model."
|
||
|
|
|
||
|
|
The library should first determine what kind of failure is happening.
|
||
|
|
|
||
|
|
Different failure types require different responses:
|
||
|
|
|
||
|
|
- some need more capacity
|
||
|
|
- some need better confidence policy
|
||
|
|
- some should be decomposed
|
||
|
|
- some should become deterministic
|
||
|
|
- some have unstable labels and should not be treated as ordinary classifiers
|
||
|
|
|
||
|
|
## Failure Classes
|
||
|
|
|
||
|
|
### 1. `capacity_limited_gate`
|
||
|
|
|
||
|
|
Description:
|
||
|
|
|
||
|
|
The gate is well-defined and reasonably stationary, but the current model tier does
|
||
|
|
not have enough capacity to separate the decision surface reliably.
|
||
|
|
|
||
|
|
Characteristic symptoms:
|
||
|
|
|
||
|
|
- performance improves with training but plateaus below target
|
||
|
|
- label boundaries appear coherent under manual inspection
|
||
|
|
- confusion is broad rather than concentrated in one structural pattern
|
||
|
|
- richer models or larger context are plausible fixes
|
||
|
|
|
||
|
|
Likely root causes:
|
||
|
|
|
||
|
|
- too little model capacity
|
||
|
|
- too little representational depth
|
||
|
|
- long-range dependencies exceed current context handling
|
||
|
|
- input compression is too lossy for the chosen tier
|
||
|
|
|
||
|
|
Recommended probes:
|
||
|
|
|
||
|
|
- larger-model comparison
|
||
|
|
- context-width probe
|
||
|
|
- calibration probe
|
||
|
|
- feature ablation probe
|
||
|
|
|
||
|
|
Recommended interventions:
|
||
|
|
|
||
|
|
- move to a larger specialist tier
|
||
|
|
- add better structured features
|
||
|
|
- use a richer encoder while preserving the same gate schema
|
||
|
|
|
||
|
|
### 2. `factorizable_gate`
|
||
|
|
|
||
|
|
Description:
|
||
|
|
|
||
|
|
The gate is represented as one multiclass decision, but the actual semantics are a
|
||
|
|
composition of smaller independent or semi-independent decisions.
|
||
|
|
|
||
|
|
Characteristic symptoms:
|
||
|
|
|
||
|
|
- flat multiclass performance despite apparently simple semantics
|
||
|
|
- classes can be expressed as Cartesian products or hierarchical branches
|
||
|
|
- confusion concentrates between composite labels
|
||
|
|
- downstream action is itself assembled from multiple sub-decisions
|
||
|
|
|
||
|
|
Likely root causes:
|
||
|
|
|
||
|
|
- a product decision has been collapsed into one label
|
||
|
|
- one model is being asked to infer multiple policy axes at once
|
||
|
|
- the current evaluation obscures easier sub-structure
|
||
|
|
|
||
|
|
Recommended probes:
|
||
|
|
|
||
|
|
- factorization probe
|
||
|
|
- hierarchical-gate probe
|
||
|
|
- per-axis confidence probe
|
||
|
|
|
||
|
|
Recommended interventions:
|
||
|
|
|
||
|
|
- split into multiple smaller gates
|
||
|
|
- add a deterministic combiner over sub-gates
|
||
|
|
- replace flat multiclass routing with hierarchical routing
|
||
|
|
|
||
|
|
### 3. `non_stationary_gate`
|
||
|
|
|
||
|
|
Description:
|
||
|
|
|
||
|
|
The label semantics change over time because they depend on evolving tooling,
|
||
|
|
policies, or environment state rather than a stable ontology.
|
||
|
|
|
||
|
|
Characteristic symptoms:
|
||
|
|
|
||
|
|
- labels become stale as the surrounding system improves
|
||
|
|
- class definitions drift without changing the input domain
|
||
|
|
- retraining is needed because the target policy moved, not because the model forgot
|
||
|
|
- disagreements cluster at capability boundaries rather than lexical boundaries
|
||
|
|
|
||
|
|
Likely root causes:
|
||
|
|
|
||
|
|
- output labels encode current deployment policy
|
||
|
|
- the gate depends on tool availability or maturity
|
||
|
|
- class meaning is roadmap-dependent
|
||
|
|
|
||
|
|
Recommended probes:
|
||
|
|
|
||
|
|
- schema-stability probe
|
||
|
|
- temporal re-label probe
|
||
|
|
- policy-variant probe
|
||
|
|
|
||
|
|
Recommended interventions:
|
||
|
|
|
||
|
|
- separate stable ontology fields from changing policy fields
|
||
|
|
- turn part of the decision into a policy layer above the model
|
||
|
|
- use hierarchical routing where the first stage predicts a stable property and a
|
||
|
|
later stage applies current deployment policy
|
||
|
|
|
||
|
|
### 4. `missing_context_gate`
|
||
|
|
|
||
|
|
Description:
|
||
|
|
|
||
|
|
The gate might be learnable, but the current input encoding omits the context needed
|
||
|
|
to make the decision safely.
|
||
|
|
|
||
|
|
Characteristic symptoms:
|
||
|
|
|
||
|
|
- short-text performance is poor but humans request surrounding context
|
||
|
|
- errors cluster on inputs that depend on hidden project or workflow state
|
||
|
|
- adding metadata is expected to help more than scaling parameters
|
||
|
|
- the same surface text would require different answers in different contexts
|
||
|
|
|
||
|
|
Likely root causes:
|
||
|
|
|
||
|
|
- important state is absent from the input
|
||
|
|
- text summaries collapse policy-critical structure
|
||
|
|
- the gate implicitly depends on project metadata, not just user language
|
||
|
|
|
||
|
|
Recommended probes:
|
||
|
|
|
||
|
|
- context-width probe
|
||
|
|
- metadata augmentation probe
|
||
|
|
- stateful-context probe
|
||
|
|
|
||
|
|
Recommended interventions:
|
||
|
|
|
||
|
|
- enrich inputs with structured metadata
|
||
|
|
- separate text understanding from stateful decision logic
|
||
|
|
- require abstention when needed context is unavailable
|
||
|
|
|
||
|
|
### 5. `deterministic_disguised_as_ml_gate`
|
||
|
|
|
||
|
|
Description:
|
||
|
|
|
||
|
|
The gate looks statistical at first glance, but the output is largely determined by
|
||
|
|
known rules or a stable formula.
|
||
|
|
|
||
|
|
Characteristic symptoms:
|
||
|
|
|
||
|
|
- very high performance with synthetic or structural encodings
|
||
|
|
- human explanation reduces to a crisp rule
|
||
|
|
- model errors mostly reflect encoding mismatch, not true ambiguity
|
||
|
|
- downstream users already trust a formula more than the classifier
|
||
|
|
|
||
|
|
Likely root causes:
|
||
|
|
|
||
|
|
- a deterministic rule was not implemented yet
|
||
|
|
- the learned model is being used as a proxy for explicit policy
|
||
|
|
- training data was easier to generate than writing the rule
|
||
|
|
|
||
|
|
Recommended probes:
|
||
|
|
|
||
|
|
- deterministic-rule probe
|
||
|
|
- symbolic baseline probe
|
||
|
|
- feature-rule extraction probe
|
||
|
|
|
||
|
|
Recommended interventions:
|
||
|
|
|
||
|
|
- replace with deterministic logic
|
||
|
|
- keep the model only as a convenience wrapper over explicit rules
|
||
|
|
- use the learned gate only when the rule inputs are partially missing
|
||
|
|
|
||
|
|
### 6. `label_space_mismatch`
|
||
|
|
|
||
|
|
Description:
|
||
|
|
|
||
|
|
The gate's labels are underspecified, overlapping, misordered, or otherwise not
|
||
|
|
aligned with the actual decision boundary the system needs.
|
||
|
|
|
||
|
|
Characteristic symptoms:
|
||
|
|
|
||
|
|
- confusion clusters around neighboring or semantically overlapping labels
|
||
|
|
- multiple labels seem equally valid for the same input
|
||
|
|
- label guidelines are hard to explain concisely
|
||
|
|
- accuracy does not reflect practical usefulness because some "mistakes" are close
|
||
|
|
|
||
|
|
Likely root causes:
|
||
|
|
|
||
|
|
- labels mix ontology and policy
|
||
|
|
- labels are too coarse or too fine
|
||
|
|
- adjacent classes are not operationally distinct
|
||
|
|
|
||
|
|
Recommended probes:
|
||
|
|
|
||
|
|
- confusion-structure probe
|
||
|
|
- class-merge probe
|
||
|
|
- hierarchical-label probe
|
||
|
|
|
||
|
|
Recommended interventions:
|
||
|
|
|
||
|
|
- merge or split labels
|
||
|
|
- add hierarchy to the label space
|
||
|
|
- redefine evaluation so high-cost confusions are distinguished from near-miss ones
|
||
|
|
|
||
|
|
### 7. `guardrail_limited_gate`
|
||
|
|
|
||
|
|
Description:
|
||
|
|
|
||
|
|
The gate is good enough to be useful, but not good enough to be trusted without a
|
||
|
|
deployment policy around it.
|
||
|
|
|
||
|
|
Characteristic symptoms:
|
||
|
|
|
||
|
|
- raw accuracy is moderate but confidence is informative
|
||
|
|
- abstention can isolate hard cases
|
||
|
|
- downstream checks can catch many residual errors
|
||
|
|
- accepted predictions are much stronger than overall predictions
|
||
|
|
|
||
|
|
Likely root causes:
|
||
|
|
|
||
|
|
- the model is solving an economically useful subset of the problem
|
||
|
|
- deployment policy has not yet been calibrated
|
||
|
|
|
||
|
|
Recommended probes:
|
||
|
|
|
||
|
|
- confidence-threshold probe
|
||
|
|
- abstain-coverage probe
|
||
|
|
- retry-recovery probe
|
||
|
|
- deterministic-check catch-rate probe
|
||
|
|
|
||
|
|
Recommended interventions:
|
||
|
|
|
||
|
|
- deploy with guardrails
|
||
|
|
- set confidence thresholds by risk tier
|
||
|
|
- add retry and escalation paths
|
||
|
|
|
||
|
|
## Diagnostic Probes
|
||
|
|
|
||
|
|
The taxonomy is only useful if the library has standard probes.
|
||
|
|
|
||
|
|
### `confidence_threshold_probe`
|
||
|
|
|
||
|
|
Question:
|
||
|
|
|
||
|
|
Does the gate become safe and useful when low-confidence cases abstain?
|
||
|
|
|
||
|
|
Outputs:
|
||
|
|
|
||
|
|
- accept rate
|
||
|
|
- accuracy on accepted
|
||
|
|
- silent error estimate
|
||
|
|
|
||
|
|
### `factorization_probe`
|
||
|
|
|
||
|
|
Question:
|
||
|
|
|
||
|
|
Can a weak multiclass gate be re-expressed as multiple smaller gates with a better
|
||
|
|
combined policy?
|
||
|
|
|
||
|
|
Outputs:
|
||
|
|
|
||
|
|
- candidate factorization structure
|
||
|
|
- estimated per-axis difficulty
|
||
|
|
- combined accuracy or accepted-precision comparison
|
||
|
|
|
||
|
|
### `deterministic_rule_probe`
|
||
|
|
|
||
|
|
Question:
|
||
|
|
|
||
|
|
Is the gate mostly encoding a crisp rule that should live in ordinary software?
|
||
|
|
|
||
|
|
Outputs:
|
||
|
|
|
||
|
|
- symbolic baseline score
|
||
|
|
- overlap between model errors and rule violations
|
||
|
|
|
||
|
|
### `schema_stability_probe`
|
||
|
|
|
||
|
|
Question:
|
||
|
|
|
||
|
|
Are the labels stable over time, or do they move with deployment policy and tool
|
||
|
|
availability?
|
||
|
|
|
||
|
|
Outputs:
|
||
|
|
|
||
|
|
- stability rating
|
||
|
|
- ontology-vs-policy split recommendation
|
||
|
|
|
||
|
|
### `context_width_probe`
|
||
|
|
|
||
|
|
Question:
|
||
|
|
|
||
|
|
Is performance limited by missing project or workflow context?
|
||
|
|
|
||
|
|
Outputs:
|
||
|
|
|
||
|
|
- delta from richer context
|
||
|
|
- recommended minimum context set
|
||
|
|
|
||
|
|
### `confusion_structure_probe`
|
||
|
|
|
||
|
|
Question:
|
||
|
|
|
||
|
|
Are the observed errors random, adjacent, hierarchical, or compositional?
|
||
|
|
|
||
|
|
Outputs:
|
||
|
|
|
||
|
|
- confusion topology
|
||
|
|
- merge, split, or hierarchy recommendation
|
||
|
|
|
||
|
|
## Recommended Response Order
|
||
|
|
|
||
|
|
The library should approach gate triage in this order:
|
||
|
|
|
||
|
|
1. Check whether the gate is deterministic in disguise.
|
||
|
|
2. Check whether the label space is stable and coherent.
|
||
|
|
3. Check whether the gate is factorizable.
|
||
|
|
4. Check whether critical context is missing.
|
||
|
|
5. Check whether a guarded deployment policy is already sufficient.
|
||
|
|
6. Only then treat scale as the primary lever.
|
||
|
|
|
||
|
|
## Library Implications
|
||
|
|
|
||
|
|
To support this taxonomy, `whetstone_RSA` should eventually represent each gate with
|
||
|
|
enough metadata for diagnosis.
|
||
|
|
|
||
|
|
Suggested gate metadata:
|
||
|
|
|
||
|
|
- `gate_id`
|
||
|
|
- `label_schema`
|
||
|
|
- `label_stability`
|
||
|
|
- `task_family`
|
||
|
|
- `input_modalities`
|
||
|
|
- `supports_abstain`
|
||
|
|
- `deterministic_baseline_available`
|
||
|
|
- `candidate_factorizations`
|
||
|
|
- `risk_tier`
|
||
|
|
|
||
|
|
Suggested diagnostic result shape:
|
||
|
|
|
||
|
|
```text
|
||
|
|
GateDiagnosis
|
||
|
|
primary_failure_class
|
||
|
|
secondary_failure_classes
|
||
|
|
supporting_signals
|
||
|
|
recommended_probes
|
||
|
|
recommended_interventions
|
||
|
|
confidence
|
||
|
|
```
|
||
|
|
|
||
|
|
## Case-Study Motivators
|
||
|
|
|
||
|
|
The case study motivates this taxonomy, but does not constrain it.
|
||
|
|
|
||
|
|
- A gate like `prereq_op` suggests `factorizable_gate`
|
||
|
|
- A gate like `automatability` suggests `non_stationary_gate`
|
||
|
|
- A gate like `confidence_tier` suggests `deterministic_disguised_as_ml_gate`
|
||
|
|
- A gate like `verification_type` or `worker_type` may become
|
||
|
|
`guardrail_limited_gate`
|
||
|
|
|
||
|
|
These examples are evidence of recurring patterns, not bespoke exceptions.
|
||
|
|
|
||
|
|
## Summary
|
||
|
|
|
||
|
|
The main architectural insight is simple:
|
||
|
|
|
||
|
|
Bounded-gate failure is not one problem.
|
||
|
|
|
||
|
|
If the RSA library can classify failure types before scaling models blindly, it will
|
||
|
|
be cheaper, safer, and easier to generalize across domains.
|