Files
whetstone_RSA/docs/diagnostic_patterns.md

435 lines
21 KiB
Markdown
Raw Normal View History

# RSA Diagnostic Patterns
## Purpose
This document is a named-pattern library for diagnosing bounded decision gates.
Each pattern has:
- a name that can be referenced in conversations and code
- a concrete observable signal
- the mechanism that produces that signal
- a detection test
- a specific fix
- the conditions under which the fix fails
- a real example from the `whetstone_DSL` case study where one exists
This is not the same as the failure taxonomy (`gate_failure_taxonomy.md`).
The taxonomy classifies failure _types_. This document gives you something to
actually do when you observe a specific symptom.
**Intended audience:** developers, model-assisted debuggers, automated diagnosis
pipelines. Each pattern is written at a level of specificity that a smaller model
or a programmer without deep ML background can apply correctly if they can
recognize the signal.
---
## Pattern Index
| Name | Signal | Fix |
|------|--------|-----|
| [Directional Collapse](#1-directional-collapse) | One class maps >80% to one other specific class | Merge those classes or route them to a prior stage |
| [Constant Class](#2-constant-class) | One label fires on 100% (or near-100%) of corpus rows | Remove label from gate; encode as a rule |
| [Corpus Uniformity Trap](#3-corpus-uniformity-trap) | High row count but very few unique input texts | Discard corpus; find a richer data source |
| [Capacity Plateau](#4-capacity-plateau) | Larger model tier does not improve over smaller | Stop scaling; diagnose label or data quality |
| [Schema Drift Contamination](#5-schema-drift-contamination) | Structured field contains out-of-distribution values from another class | Hard-filter before training |
| [Upstream Feature Starvation](#6-upstream-feature-starvation) | Gate performance doesn't improve with any intervention | Check whether discriminative fields exist at inference time |
| [Axis Factorization](#7-axis-factorization) | Multi-class gate is weak; classes can be expressed as a Cartesian product | Decompose into binary sub-gates with a deterministic combiner |
| [Grokking Phase Transition](#8-grokking-phase-transition) | Accuracy is flat for thousands of steps, then jumps sharply | Do not stop training at the plateau |
| [Guardrail Rescue](#9-guardrail-rescue) | Raw accuracy is moderate but confidence is informative | Deploy with an abstain threshold rather than scaling |
| [Label Boundary Drift](#10-label-boundary-drift) | Confusion is broad and inconsistent; manual review disagrees on labeling | Redefine labels, not model |
---
## 1. Directional Collapse
**Signal:** In the confusion matrix, one class maps overwhelmingly (>80%) to one
other specific class. The confusion is directional and concentrated, not diffuse.
**Mechanism:** The two classes are indistinguishable in the input representation.
The model has learned everything it can and is guessing the majority direction.
Scaling or adding data will not help unless you change what features the gate sees.
This is different from random confusion (which is diffuse) and from capacity
limitation (which produces broad, scattered errors).
**Detection test:** Look at the per-class confusion. If class A sends >80% of its
errors to exactly one destination class B, and this persists across model tiers,
you have Directional Collapse, not a capacity problem.
**Fix:** Merge A and B into a single bucket. If the downstream system needs to
distinguish them, add a prior routing stage that can separate them from the rest
of the label space — then apply a second gate only to that bucket. The two-stage
decomposition is the standard fix.
**When the fix fails:** If the merged bucket is large and the downstream system
genuinely needs the A/B distinction, you have a feature enrichment problem, not a
model problem. The fix is to add a feature to the input schema that carries the
A/B signal, then retrain.
**Taxonomy link:** `label_space_mismatch` in `gate_failure_taxonomy.md`.
**Example — `automatability` gate, April 2026:**
The `deterministic` class sent 100% of its errors to `template` in the medium-tier
model. Adding data and scaling both failed. The fix was a two-stage decomposition:
stage 1 (binary) separates `{deterministic, template}` from `{specialist, slm, llm,
human}`; stage 2 (4-way) classifies within the routed set. Both stages achieved 100%
accuracy. The `deterministic`/`template` distinction turns out to require information
not present in taskitem titles — that distinction belongs in a richer input schema,
not in the current gate.
---
## 2. Constant Class
**Signal:** One label in your label schema fires on 100% (or near-100%) of all rows
in the corpus. It is always true regardless of input.
**Mechanism:** The label is not a decision surface — it is a structural property of
the corpus or the system. Every example in this data source satisfies the condition,
so no model can learn when it _doesn't_ fire.
**Detection test:** Compute per-label frequency over the full corpus. If any label
is present in ≥95% of rows, apply this pattern.
**Fix:** Remove the label from the gate's label space. If the condition is always
true in the current system, encode it as a rule (`always apply X`). If it may
eventually become false as the system evolves, defer the gate until real negative
examples exist.
**When the fix fails:** If the label was constant only in the current corpus and
will have genuine variance in a future corpus (different data sources, different
system states), removing it may be premature. Verify by asking whether a different
project or data source would produce negative examples.
**Taxonomy link:** `deterministic_disguised_as_ml_gate` in `gate_failure_taxonomy.md`.
**Example — `needs_validate_intake` field, April 2026:**
Every one of 1,297 accepted rows had `needs_validate_intake = True`. The field was
in the label space because it appeared to be a decision output, but inspection
showed it was always True in the corpus. It was dropped from the label space. The
effective gate became two binary classifiers over the remaining fields.
---
## 3. Corpus Uniformity Trap
**Signal:** The corpus has high row count but very few unique input texts. For
example: 1,297 rows with only 17 unique texts.
**Mechanism:** The data source generates many rows from a small number of canonical
templates. The apparent training set size is misleading — the model is effectively
seeing a handful of inputs repeated. Variance in outputs cannot be attributed to
variance in inputs, so the model cannot learn a discriminative function.
**Detection test:** Count unique input texts. If `unique_texts / total_rows < 0.1`,
or if unique text count is below ~50 regardless of total rows, you are in this trap.
**Fix:** Find a different data source that has genuine input variance. Do not
generate more rows from the same corpus. Do not oversample or augment the existing
rows. The problem is the source, not the quantity.
**When the fix fails:** If the only available data is template-driven (e.g., all
examples come from the same intake spec format), this pattern is a blocking
constraint on the gate. The gate cannot be trained until the input schema or data
source provides discriminative variance. Document this as a `missing_context_gate`
and treat data collection as a prerequisite milestone.
**Taxonomy link:** `missing_context_gate` in `gate_failure_taxonomy.md`.
**Example — `prereq_op` gate, April 2026:**
The whetstone_DSL run corpus produced 1,297 rows but only 17 unique context texts.
All tasks generated from the same intake spec had identical context text. The
discriminative signal (step-level description) was never stored in the run artifact.
Resolution: add RSA-relevant fields to the taskitem schema so that each task is
self-contained. Collection of real training data is blocked until that schema change.
---
## 4. Capacity Plateau
**Signal:** A larger model tier (e.g., medium vs. small_plus) does not improve
accuracy over the smaller tier, or slightly regresses. The performance gap between
tiers is within noise.
**Mechanism:** The bottleneck is not model capacity — it is data quality, label
noise, or feature representation. More parameters cannot overcome bad labels or
indistinguishable input features. The model has already saturated what the data
supports.
**Detection test:** If tier comparison shows improvement < 2pp from small_plus to
medium, and confusion patterns are similar across tiers, apply this pattern.
**Fix:** Stop scaling. Diagnose label quality, input feature richness, and corpus
diversity instead. In priority order:
1. Check for Directional Collapse (Pattern 1)
2. Check for Corpus Uniformity Trap (Pattern 3)
3. Check for Upstream Feature Starvation (Pattern 6)
4. Check for Axis Factorization (Pattern 7)
**When the fix fails:** If none of the above patterns apply and the gate genuinely
requires more representational power, scale is the right lever. But this should be
the last resort, not the first response.
**Taxonomy link:** Capacity plateau is a meta-signal that rules out
`capacity_limited_gate` and shifts suspicion to other failure classes.
**Example — `worker_type` and `automatability` gates, April 2026:**
Both gates showed near-zero improvement (and slight regression) from small_plus to
medium. For `automatability`, the root cause was Directional Collapse. For
`worker_type`, the root cause is label noise from heuristic classification — the
heuristic assigns ambiguous titles based on keyword matching, producing inconsistent
boundaries at training time.
---
## 5. Schema Drift Contamination
**Signal:** A structured output field contains values that belong to a different
schema or a different version of the contract. For example, a field that should
contain `{standard, needs_review, needs_approval, full_gates}` instead contains
`whetstone_generate_taskitems` or `manual-approval` in a subset of rows.
**Mechanism:** The upstream system that populated the field was in a different mode
(fallback behavior, older version, or a tool that doesn't respect the current
contract). The contamination rows are not wrong in the training data sense — they
are accurately labeled artifacts of a real system state — but they represent a
different underlying distribution.
**Detection test:** For each structured field, enumerate all unique values and check
for out-of-vocabulary entries against the current label schema.
**Fix:** Hard-filter contaminated rows before training. Do not model drift as a
separate class unless you are deliberately building a drift detector. If you want
to handle drift gracefully at inference time, build a separate pre-filter gate.
**When the fix fails:** If the fraction of contaminated rows is large (>20%), the
root cause is likely a systemic upstream problem. Filtering removes the symptom but
not the cause. Document the contamination source and fix the upstream system.
**Taxonomy link:** `non_stationary_gate` in `gate_failure_taxonomy.md`.
**Example — `prereq_op` extraction, April 2026:**
Two contamination patterns were found in the extracted corpus: tool-ID drift
(12 rows with `whetstone_generate_taskitems` etc. in `prerequisiteOps`) and
`manual-approval` drift (12 rows with an undocumented op name). Both were
hard-filtered. The extraction tool correctly flagged these via rejection logic.
---
## 6. Upstream Feature Starvation
**Signal:** Gate accuracy does not improve regardless of model tier, data volume,
or label quality interventions. The gate looks solvable in principle but resists
all fixes.
**Mechanism:** The information needed to make the decision correctly does not exist
in the input schema at inference time. The model is not failing to learn — it is
being asked to infer a signal that was never provided. This is not a model problem.
**Detection test:** Ask: "If a human were shown only what the model sees at
inference time, could the human reliably make this decision?" If the answer is no,
this pattern applies. Also: check whether discriminative fields are present in
training data sources but absent from the inference-time schema.
**Fix:** Enrich the upstream schema so the discriminative fields are available when
the gate runs. This is a schema change, not a model change. Retraining after the
schema change will likely produce rapid improvement.
**When the fix fails:** If the upstream system cannot practically carry the required
fields (e.g., they depend on information only available later in the pipeline), the
gate may need to move downstream to where the information exists, or abstain and
escalate until context is available.
**Taxonomy link:** `missing_context_gate` in `gate_failure_taxonomy.md`.
**Example — `prereq_op` and `automatability` gates, April 2026:**
The RSA analysis found that the taskitem schema does not carry fields like
`architectural_surface`, `cross_component_deps`, or `security_sensitive` at
inference time. These are the discriminative features for several gates.
`generate_taskitems` makes 10+ separable decisions in one LLM call; the RSA
layer cannot see the intermediate reasoning. Conclusion: the taskitem schema must
carry RSA-relevant fields as first-class outputs of `generate_taskitems` before
reliable training data can be collected for most gates.
---
## 7. Axis Factorization
**Signal:** A multi-class gate has weak accuracy despite apparently simple semantics.
The classes can be described as a Cartesian product or hierarchical branching of two
or more independent binary conditions.
**Mechanism:** The gate is being asked to make multiple decisions at once with a
single output label. The flat label space hides compositional structure. A model
cannot reliably learn a joint decision that is actually two independent decisions
mixed together.
**Detection test:** Try to describe each class as a combination of simpler
conditions. If you can express the full label set as `condition_A × condition_B`,
factorization applies. Also look for confusion that concentrates between classes
that share one axis but differ on the other.
**Fix:** Decompose into one binary gate per axis. Add a deterministic combiner that
maps the binary outputs back to the original multi-class label. Evaluate the
combined accuracy against the flat gate.
**When the fix fails:** If the axes are not independent (the output of one
sub-decision changes the valid space of the other), full factorization is not safe.
In that case, use hierarchical routing instead: a first gate routes to a subset,
a second gate classifies within that subset.
**Taxonomy link:** `factorizable_gate` in `gate_failure_taxonomy.md`.
**Example — `prereq_op` gate, April 2026:**
The four classes (`standard`, `needs_review`, `needs_approval`, `full_gates`) are
the Cartesian product of two binary axes: `requires_architect_review` and
`requires_manual_approval`. Decomposing into two binary classifiers and combining
deterministically gives combined accuracy directly comparable to the flat gate,
with the added benefit that each sub-gate is independently trainable, debuggable,
and deployable.
---
## 8. Grokking Phase Transition
**Signal:** Accuracy is flat or near-flat for a long training run, then jumps
sharply by 1020pp in a short window. The jump is sustained — the model does not
regress after the transition.
**Mechanism:** The model is slowly building internal representations before they
cross a threshold of usefulness. This is a known phenomenon in small transformers
trained on structured tasks. The plateau is not stagnation — it is slow
generalization development. Stopping during the plateau will miss the transition.
**Detection test:** If accuracy is stagnant for thousands of steps but loss is
still decreasing, you may be approaching a transition. A sudden jump in accuracy
with sustained improvement confirms this pattern.
**Fix:** Do not stop training at the plateau. Set a patience window or use loss
as the primary stopping criterion rather than accuracy. Run long enough to confirm
whether a transition occurs.
**When the fix fails:** If loss stops decreasing and accuracy remains flat, the
plateau is genuine convergence on a suboptimal local minimum — not a pre-transition
state. In that case, grokking is not occurring and you should investigate data or
architecture.
**Taxonomy link:** Not a failure pattern — a learning dynamic. May initially look
like `capacity_limited_gate` but is resolved by training duration, not model size.
---
## 9. Guardrail Rescue
**Signal:** Raw accuracy is moderate (6080%) but threshold analysis shows that
accuracy on high-confidence predictions is substantially higher (>90%), with
acceptable abstain rates at the required threshold.
**Mechanism:** The gate is solving a tractable subset of the problem confidently
and correctly. The remaining cases are genuinely hard — not model failures, but
irreducible uncertainty from the input. A deployment policy that abstains on low
confidence provides real value even without improving the model.
**Detection test:** Check threshold analysis: `accuracy_on_accepted` at `threshold=0.9`
vs. `raw_accuracy`. If the gap is >10pp with `abstain_rate < 0.30`, Guardrail Rescue
applies. The gate is usable today.
**Fix:** Deploy with an abstain threshold calibrated to the gate's risk tier. Route
abstentions to a retry (with enriched context) or escalation path. Do not block
deployment on achieving target raw accuracy if the guardrail profile is acceptable.
**When the fix fails:** If high-confidence predictions are not substantially better
than raw accuracy (confidence is not calibrated or informative), the model is
overconfident and guardrail deployment is not safe. Check calibration before
deploying with guardrails.
**Taxonomy link:** `guardrail_limited_gate` in `gate_failure_taxonomy.md`.
**Example — `automatability` small_plus, April 2026:**
Raw accuracy was 69.6%, but at threshold=0.9, accuracy on accepted predictions
was 100% — with 30.4% abstain rate. The gate correctly handles 70% of inputs
with perfect precision. The abstentions are the genuinely hard cases (primarily
`deterministic` and `template` boundary). This profile was superseded by the
two-stage decomposition, but the guardrail analysis would have been valid as an
intermediate deployment.
---
## 10. Label Boundary Drift
**Signal:** Confusion is broad and scattered — no single directional pattern
dominates. Manual review of errors produces disagreement between reviewers about
the correct label. Retraining on more data does not converge.
**Mechanism:** The label definitions are ambiguous or inconsistent. The model is
learning to imitate noisy, inconsistent labels. It is performing correctly against
the training signal it has — the training signal is the problem.
**Detection test:** Manually review 2030 errors. If reviewers disagree on the
correct label for more than 30% of examples, Label Boundary Drift applies.
Also: if heuristic-labeled data and human-labeled data disagree in systematic ways,
this is the pattern.
**Fix:** Rewrite the label definitions. Add concrete decision criteria that resolve
ambiguous cases. Relabel a sample and check inter-annotator agreement before
retraining. This is a labeling problem, not a model problem.
**When the fix fails:** If the boundary is genuinely contextual (the same input
requires different labels in different system states), the labels may need to
incorporate system state as a feature. See also Upstream Feature Starvation
(Pattern 6).
**Taxonomy link:** `label_space_mismatch` in `gate_failure_taxonomy.md`.
**Example — `worker_type` gate, ongoing:**
The gate uses heuristic keyword classification to generate labels. Tasks like
"Snapshot persistence (sidecar integration)" or "Compliance policy binding engine"
are labeled `implementer` by the heuristic but the model predicts `architect` or
`reviewer`. Both predictions are arguable from the title alone. The heuristic
assigns based on keyword absence (implementer is the default), which produces
unreliable labels at the boundary. Resolution requires either human labeling of
ambiguous cases or richer input features that remove the ambiguity.
---
## How To Use This Document
When a gate is underperforming, check the patterns in this order:
1. **Look at the confusion matrix.** Is any class sending >80% of errors to one
specific other class? → [Directional Collapse](#1-directional-collapse)
2. **Look at label frequency.** Is any label firing on >95% of corpus rows?
→ [Constant Class](#2-constant-class)
3. **Count unique input texts.** Is `unique_texts / total_rows < 0.1`?
→ [Corpus Uniformity Trap](#3-corpus-uniformity-trap)
4. **Compare tiers.** Did the larger model fail to improve?
→ [Capacity Plateau](#4-capacity-plateau), then diagnose further
5. **Check the input schema.** Does the information needed for the decision
actually exist in the input at inference time?
→ [Upstream Feature Starvation](#6-upstream-feature-starvation)
6. **Check structured fields.** Do any fields contain out-of-vocabulary values?
→ [Schema Drift Contamination](#5-schema-drift-contamination)
7. **Try to describe the classes as a Cartesian product.**
→ [Axis Factorization](#7-axis-factorization)
8. **Check threshold analysis.** Does high-confidence accuracy substantially
exceed raw accuracy?
→ [Guardrail Rescue](#9-guardrail-rescue)
9. **Manually review errors.** Do reviewers disagree on the labels?
→ [Label Boundary Drift](#10-label-boundary-drift)
10. **Check training curves.** Was the model stopped during a long plateau before
a possible grokking transition?
→ [Grokking Phase Transition](#8-grokking-phase-transition)
Scale (larger model, more data) is not a pattern. It is the default assumption
that these patterns exist to refute.