5.4 KiB
Sprint 001: Gate Failure Taxonomy
Status
In Progress
Intent
Define an abstract, library-level framework for diagnosing why bounded decision
gates fail and what restructuring options whetstone_RSA should support.
This sprint is explicitly about the RSA library.
It is not about changing whetstone_DSL.
whetstone_DSL is used only as a case study to surface recurring gate failure
patterns that the library should handle generically.
Problem
Some bounded gates train well at tiny model size. Others plateau at mediocre accuracy, and it is unclear whether the right response is:
- a larger model
- a better confidence policy
- a decomposed gate structure
- a deterministic pre-pass
- richer input context
- or a relabeling of the decision space
Right now that diagnosis depends too much on a knowledgeable human noticing the pattern by inspection.
The RSA library should instead provide abstract mechanisms for:
- identifying failure signatures
- classifying likely root causes
- probing alternative gate structures
- recommending safer deployment policies
Scope
This sprint produces a first-pass taxonomy and design contract for gate diagnosis.
Primary output:
docs/gate_failure_taxonomy.md
Supporting updates as needed:
docs/architecture.mddocs/evaluation_policy.mddocs/gate_rubric.md
Case Study Constraint
Use whetstone_DSL/specialists as evidence only.
Do not propose direct code changes to whetstone_DSL as the sprint output.
Instead, convert the observed patterns into reusable RSA abstractions.
Working Hypothesis
Poor bounded-gate performance can usually be grouped into a small number of failure types that should be recognized by the library:
capacity_limited_gatefactorizable_gatenon_stationary_gatemissing_context_gatedeterministic_disguised_as_ml_gatelabel_space_mismatch
The library should be able to recommend different responses for each failure type without requiring a human to rediscover the pattern each time.
Deliverables
1. Failure Taxonomy
Define each failure class with:
- description
- characteristic symptoms
- likely root causes
- recommended probes
- recommended interventions
2. Probe Model
Define library-level diagnostic probes such as:
- confidence-threshold probe
- class-confusion probe
- factorization probe
- deterministic-rule probe
- schema-stability probe
- context-width probe
These are conceptual probes first, not necessarily implemented code in this sprint.
3. Restructuring Policy
Define abstract restructuring options available to RSA:
- keep as single multiclass gate
- factor into multiple smaller gates
- convert part of the problem into deterministic rules
- change to hierarchical routing
- abstain and escalate
- relabel or merge classes
4. Library Contract Implications
Describe what the runtime should eventually support so this taxonomy is usable:
- gate metadata sufficient for diagnosis
- pluggable probes
- structured recommendations
- logging fields that make diagnosis possible
Why This Matters
Without a failure taxonomy, model-size escalation becomes the default answer to any weak gate.
That is the wrong default.
The library should first ask:
- is this actually one gate?
- is the label space stable?
- is the model missing critical context?
- is part of this already deterministic?
- is the policy target changing as the system evolves?
Only after those questions are answered should model scale be the main lever.
WhetstoneDSL-Derived Evidence To Capture
prereq_op
Observed pattern:
- tiny-tier performance around
69% - combined 4-way label built from two latent binary decisions
- heuristic keyword labeling over short text
Library lesson:
- multiclass gates may hide compositional structure
- factorization probes should be first-class
automatability
Observed pattern:
- tiny-tier performance around
69% - labels depend partly on current system capabilities and tooling maturity
Library lesson:
- some gates are non-stationary
- schema-stability checks must be first-class
confidence_tier
Observed pattern:
- perfect or near-perfect behavior because the task is effectively formulaic
Library lesson:
- the library should detect when a gate belongs in deterministic tooling rather than learned routing
Acceptance Criteria
- A new taxonomy document exists in
whetstone_RSA/docs/ - The taxonomy is framed in RSA-library terms, not WhetstoneDSL implementation terms
- The document includes at least five distinct failure classes
- Each class includes symptoms, probes, and recommended interventions
- The case study evidence is cited only as motivating evidence
- The sprint file is sufficient for a later session to resume work cleanly
Not In Scope
- modifying
whetstone_DSL - retraining any current specialist
- implementing all probes in code
- deciding final thresholds for every future application
Completed In This Sprint
- Wrote
docs/gate_failure_taxonomy.md - Updated
docs/architecture.mdto include gate diagnosis as a first-class concern - Updated
docs/evaluation_policy.mdto route weak-gate triage through the taxonomy - Wrote
docs/gate_rubric.mdto translate the case study into provisional policy
Next Step
Define the diagnosis API and probe contracts so the taxonomy can become executable library behavior rather than remaining documentation only.
Target artifact:
docs/diagnosis_api.md