291 lines
6.2 KiB
Markdown
291 lines
6.2 KiB
Markdown
|
|
# Probe Execution Flow
|
|||
|
|
|
|||
|
|
## Purpose
|
|||
|
|
|
|||
|
|
This document describes how `whetstone_RSA` should execute gate diagnosis in
|
|||
|
|
practice.
|
|||
|
|
|
|||
|
|
The taxonomy and diagnosis API define what the library needs to represent.
|
|||
|
|
This document defines how the pieces fit together during a diagnosis run.
|
|||
|
|
|
|||
|
|
## High-Level Flow
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
GateDefinition
|
|||
|
|
+
|
|||
|
|
GateEvidence
|
|||
|
|
+
|
|||
|
|
PolicyContext
|
|||
|
|
↓
|
|||
|
|
Probe selection
|
|||
|
|
↓
|
|||
|
|
Applicable probes run
|
|||
|
|
↓
|
|||
|
|
Probe results normalized
|
|||
|
|
↓
|
|||
|
|
Failure-class synthesis
|
|||
|
|
↓
|
|||
|
|
Intervention ranking
|
|||
|
|
↓
|
|||
|
|
Policy recommendation
|
|||
|
|
↓
|
|||
|
|
GateDiagnosis
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Inputs
|
|||
|
|
|
|||
|
|
### 1. `GateDefinition`
|
|||
|
|
|
|||
|
|
Static metadata about the gate:
|
|||
|
|
|
|||
|
|
- labels
|
|||
|
|
- slots
|
|||
|
|
- constraints
|
|||
|
|
- risk tier
|
|||
|
|
- supports abstain
|
|||
|
|
- deterministic baseline availability
|
|||
|
|
- label stability
|
|||
|
|
- candidate factorizations
|
|||
|
|
|
|||
|
|
### 2. `GateEvidence`
|
|||
|
|
|
|||
|
|
Observed evidence for one gate and one model/policy configuration:
|
|||
|
|
|
|||
|
|
- raw accuracy
|
|||
|
|
- accepted accuracy
|
|||
|
|
- abstain rate
|
|||
|
|
- retry recovery
|
|||
|
|
- silent error rate
|
|||
|
|
- calibration
|
|||
|
|
- confusion summary
|
|||
|
|
- latency
|
|||
|
|
- cost
|
|||
|
|
|
|||
|
|
### 3. `PolicyContext`
|
|||
|
|
|
|||
|
|
Deployment context that affects recommendations:
|
|||
|
|
|
|||
|
|
- risk budget
|
|||
|
|
- acceptable latency
|
|||
|
|
- acceptable compute cost
|
|||
|
|
- available escalation targets
|
|||
|
|
- whether deterministic checks exist downstream
|
|||
|
|
|
|||
|
|
## Step 1: Validate Evidence
|
|||
|
|
|
|||
|
|
Before diagnosis starts, the library should verify the evidence is usable.
|
|||
|
|
|
|||
|
|
Checks:
|
|||
|
|
|
|||
|
|
- gate id matches between definition and evidence
|
|||
|
|
- labels in confusion summary match the schema
|
|||
|
|
- metrics are internally consistent
|
|||
|
|
- risk tier and policy context are present
|
|||
|
|
|
|||
|
|
If evidence is incomplete, diagnosis should still proceed when possible, but the
|
|||
|
|
missing fields should reduce confidence in the result.
|
|||
|
|
|
|||
|
|
## Step 2: Select Applicable Probes
|
|||
|
|
|
|||
|
|
Not every probe applies to every gate.
|
|||
|
|
|
|||
|
|
Examples:
|
|||
|
|
|
|||
|
|
- `deterministic_rule_probe` applies only when structured features or a deterministic
|
|||
|
|
baseline are available
|
|||
|
|
- `factorization_probe` applies when labels suggest composition or a candidate
|
|||
|
|
factorization is registered
|
|||
|
|
- `schema_stability_probe` applies when labels may drift with policy or capability
|
|||
|
|
- `confidence_threshold_probe` applies only if confidence-bearing outputs exist
|
|||
|
|
|
|||
|
|
The library should ask each probe:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
supports(gate_definition, evidence) -> bool
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
and only run probes that opt in.
|
|||
|
|
|
|||
|
|
## Step 3: Run Probes
|
|||
|
|
|
|||
|
|
Each applicable probe returns:
|
|||
|
|
|
|||
|
|
- probe-specific signals
|
|||
|
|
- recommendation hints
|
|||
|
|
- a confidence score
|
|||
|
|
|
|||
|
|
Probe outputs should be normalized into a standard envelope before synthesis.
|
|||
|
|
|
|||
|
|
Example normalized structure:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
ProbeResult
|
|||
|
|
probe_id
|
|||
|
|
status
|
|||
|
|
signals
|
|||
|
|
recommendation_hints
|
|||
|
|
confidence
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Step 4: Aggregate Signals
|
|||
|
|
|
|||
|
|
The library should not map one probe directly to one final diagnosis.
|
|||
|
|
|
|||
|
|
Instead it should synthesize across probe outputs.
|
|||
|
|
|
|||
|
|
Examples:
|
|||
|
|
|
|||
|
|
- strong `factorization_probe` + compositional confusion pattern
|
|||
|
|
-> `factorizable_gate`
|
|||
|
|
- strong `schema_stability_probe` + drifting label semantics
|
|||
|
|
-> `non_stationary_gate`
|
|||
|
|
- strong `confidence_threshold_probe` + weak raw accuracy but high accepted precision
|
|||
|
|
-> `guardrail_limited_gate`
|
|||
|
|
- strong `deterministic_rule_probe`
|
|||
|
|
-> `deterministic_disguised_as_ml_gate`
|
|||
|
|
|
|||
|
|
This step should produce:
|
|||
|
|
|
|||
|
|
- ranked failure classes
|
|||
|
|
- supporting signals
|
|||
|
|
- unresolved ambiguities
|
|||
|
|
|
|||
|
|
## Step 5: Rank Interventions
|
|||
|
|
|
|||
|
|
Once failure classes are ranked, the library should rank interventions.
|
|||
|
|
|
|||
|
|
Interventions should come from a controlled vocabulary:
|
|||
|
|
|
|||
|
|
- `keep_current_gate`
|
|||
|
|
- `deploy_with_guardrails`
|
|||
|
|
- `raise_confidence_threshold`
|
|||
|
|
- `add_retry_with_context`
|
|||
|
|
- `add_deterministic_prepass`
|
|||
|
|
- `replace_with_deterministic_logic`
|
|||
|
|
- `factor_into_subgates`
|
|||
|
|
- `switch_to_hierarchical_gate`
|
|||
|
|
- `revise_label_schema`
|
|||
|
|
- `enrich_input_context`
|
|||
|
|
- `evaluate_larger_model_tier`
|
|||
|
|
|
|||
|
|
The order matters.
|
|||
|
|
|
|||
|
|
The library should prefer lower-cost, structurally-corrective interventions before
|
|||
|
|
defaulting to bigger models.
|
|||
|
|
|
|||
|
|
## Step 6: Emit Policy Recommendation
|
|||
|
|
|
|||
|
|
The final diagnosis should include a recommended deployment posture.
|
|||
|
|
|
|||
|
|
Possible deployment modes:
|
|||
|
|
|
|||
|
|
- `auto_accept`
|
|||
|
|
- `guarded_accept`
|
|||
|
|
- `abstain_first`
|
|||
|
|
- `research_only`
|
|||
|
|
|
|||
|
|
This recommendation should consider:
|
|||
|
|
|
|||
|
|
- risk tier
|
|||
|
|
- silent error estimate
|
|||
|
|
- accepted precision
|
|||
|
|
- latency budget
|
|||
|
|
- escalation availability
|
|||
|
|
|
|||
|
|
## Example Flow Patterns
|
|||
|
|
|
|||
|
|
### Pattern A: Good Gate With Policy Gaps
|
|||
|
|
|
|||
|
|
Observed:
|
|||
|
|
|
|||
|
|
- moderate raw accuracy
|
|||
|
|
- strong accepted precision under thresholding
|
|||
|
|
- downstream deterministic checks exist
|
|||
|
|
|
|||
|
|
Likely result:
|
|||
|
|
|
|||
|
|
- primary failure class: `guardrail_limited_gate`
|
|||
|
|
- intervention: `deploy_with_guardrails`
|
|||
|
|
|
|||
|
|
### Pattern B: Flat Multiclass Gate
|
|||
|
|
|
|||
|
|
Observed:
|
|||
|
|
|
|||
|
|
- low raw accuracy
|
|||
|
|
- confusion concentrates between composite labels
|
|||
|
|
- factorization metadata exists
|
|||
|
|
|
|||
|
|
Likely result:
|
|||
|
|
|
|||
|
|
- primary failure class: `factorizable_gate`
|
|||
|
|
- intervention: `factor_into_subgates`
|
|||
|
|
|
|||
|
|
### Pattern C: Formula-Like Gate
|
|||
|
|
|
|||
|
|
Observed:
|
|||
|
|
|
|||
|
|
- high score under structural encoding
|
|||
|
|
- symbolic baseline exists
|
|||
|
|
- class meaning is crisp
|
|||
|
|
|
|||
|
|
Likely result:
|
|||
|
|
|
|||
|
|
- primary failure class: `deterministic_disguised_as_ml_gate`
|
|||
|
|
- intervention: `replace_with_deterministic_logic`
|
|||
|
|
|
|||
|
|
### Pattern D: Drifting Capability Boundary
|
|||
|
|
|
|||
|
|
Observed:
|
|||
|
|
|
|||
|
|
- disagreements cluster between adjacent routing tiers
|
|||
|
|
- label meanings depend on current tooling maturity
|
|||
|
|
|
|||
|
|
Likely result:
|
|||
|
|
|
|||
|
|
- primary failure class: `non_stationary_gate`
|
|||
|
|
- intervention: `revise_label_schema` or separate ontology from policy
|
|||
|
|
|
|||
|
|
## Confidence Model
|
|||
|
|
|
|||
|
|
The diagnosis should carry its own confidence.
|
|||
|
|
|
|||
|
|
Diagnosis confidence should depend on:
|
|||
|
|
|
|||
|
|
- probe coverage
|
|||
|
|
- evidence completeness
|
|||
|
|
- signal agreement across probes
|
|||
|
|
- stability of the ranked recommendation
|
|||
|
|
|
|||
|
|
Low-confidence diagnosis is still useful if it says:
|
|||
|
|
|
|||
|
|
- evidence insufficient
|
|||
|
|
- collect richer context
|
|||
|
|
- run a missing probe
|
|||
|
|
|
|||
|
|
## Failure Handling
|
|||
|
|
|
|||
|
|
Diagnosis should fail soft.
|
|||
|
|
|
|||
|
|
If probes cannot reach a strong conclusion, the output should still be structured:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
GateDiagnosis
|
|||
|
|
primary_failure_class = unknown
|
|||
|
|
recommended_probes = [...]
|
|||
|
|
recommended_interventions = [collect_more_evidence]
|
|||
|
|
confidence = low
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
That is better than forcing a false diagnosis.
|
|||
|
|
|
|||
|
|
## Why This Matters
|
|||
|
|
|
|||
|
|
This execution flow is what keeps `whetstone_RSA` from becoming just a model zoo.
|
|||
|
|
|
|||
|
|
The library’s value is not only that it can run tiny models. It is that it can
|
|||
|
|
help decide:
|
|||
|
|
|
|||
|
|
- whether a gate is worth scaling
|
|||
|
|
- whether a gate should be restructured
|
|||
|
|
- whether the problem should leave ML entirely
|