6.2 KiB
Probe Execution Flow
Purpose
This document describes how whetstone_RSA should execute gate diagnosis in
practice.
The taxonomy and diagnosis API define what the library needs to represent. This document defines how the pieces fit together during a diagnosis run.
High-Level Flow
GateDefinition
+
GateEvidence
+
PolicyContext
↓
Probe selection
↓
Applicable probes run
↓
Probe results normalized
↓
Failure-class synthesis
↓
Intervention ranking
↓
Policy recommendation
↓
GateDiagnosis
Inputs
1. GateDefinition
Static metadata about the gate:
- labels
- slots
- constraints
- risk tier
- supports abstain
- deterministic baseline availability
- label stability
- candidate factorizations
2. GateEvidence
Observed evidence for one gate and one model/policy configuration:
- raw accuracy
- accepted accuracy
- abstain rate
- retry recovery
- silent error rate
- calibration
- confusion summary
- latency
- cost
3. PolicyContext
Deployment context that affects recommendations:
- risk budget
- acceptable latency
- acceptable compute cost
- available escalation targets
- whether deterministic checks exist downstream
Step 1: Validate Evidence
Before diagnosis starts, the library should verify the evidence is usable.
Checks:
- gate id matches between definition and evidence
- labels in confusion summary match the schema
- metrics are internally consistent
- risk tier and policy context are present
If evidence is incomplete, diagnosis should still proceed when possible, but the missing fields should reduce confidence in the result.
Step 2: Select Applicable Probes
Not every probe applies to every gate.
Examples:
deterministic_rule_probeapplies only when structured features or a deterministic baseline are availablefactorization_probeapplies when labels suggest composition or a candidate factorization is registeredschema_stability_probeapplies when labels may drift with policy or capabilityconfidence_threshold_probeapplies only if confidence-bearing outputs exist
The library should ask each probe:
supports(gate_definition, evidence) -> bool
and only run probes that opt in.
Step 3: Run Probes
Each applicable probe returns:
- probe-specific signals
- recommendation hints
- a confidence score
Probe outputs should be normalized into a standard envelope before synthesis.
Example normalized structure:
ProbeResult
probe_id
status
signals
recommendation_hints
confidence
Step 4: Aggregate Signals
The library should not map one probe directly to one final diagnosis.
Instead it should synthesize across probe outputs.
Examples:
- strong
factorization_probe+ compositional confusion pattern ->factorizable_gate - strong
schema_stability_probe+ drifting label semantics ->non_stationary_gate - strong
confidence_threshold_probe+ weak raw accuracy but high accepted precision ->guardrail_limited_gate - strong
deterministic_rule_probe->deterministic_disguised_as_ml_gate
This step should produce:
- ranked failure classes
- supporting signals
- unresolved ambiguities
Step 5: Rank Interventions
Once failure classes are ranked, the library should rank interventions.
Interventions should come from a controlled vocabulary:
keep_current_gatedeploy_with_guardrailsraise_confidence_thresholdadd_retry_with_contextadd_deterministic_prepassreplace_with_deterministic_logicfactor_into_subgatesswitch_to_hierarchical_gaterevise_label_schemaenrich_input_contextevaluate_larger_model_tier
The order matters.
The library should prefer lower-cost, structurally-corrective interventions before defaulting to bigger models.
Step 6: Emit Policy Recommendation
The final diagnosis should include a recommended deployment posture.
Possible deployment modes:
auto_acceptguarded_acceptabstain_firstresearch_only
This recommendation should consider:
- risk tier
- silent error estimate
- accepted precision
- latency budget
- escalation availability
Example Flow Patterns
Pattern A: Good Gate With Policy Gaps
Observed:
- moderate raw accuracy
- strong accepted precision under thresholding
- downstream deterministic checks exist
Likely result:
- primary failure class:
guardrail_limited_gate - intervention:
deploy_with_guardrails
Pattern B: Flat Multiclass Gate
Observed:
- low raw accuracy
- confusion concentrates between composite labels
- factorization metadata exists
Likely result:
- primary failure class:
factorizable_gate - intervention:
factor_into_subgates
Pattern C: Formula-Like Gate
Observed:
- high score under structural encoding
- symbolic baseline exists
- class meaning is crisp
Likely result:
- primary failure class:
deterministic_disguised_as_ml_gate - intervention:
replace_with_deterministic_logic
Pattern D: Drifting Capability Boundary
Observed:
- disagreements cluster between adjacent routing tiers
- label meanings depend on current tooling maturity
Likely result:
- primary failure class:
non_stationary_gate - intervention:
revise_label_schemaor separate ontology from policy
Confidence Model
The diagnosis should carry its own confidence.
Diagnosis confidence should depend on:
- probe coverage
- evidence completeness
- signal agreement across probes
- stability of the ranked recommendation
Low-confidence diagnosis is still useful if it says:
- evidence insufficient
- collect richer context
- run a missing probe
Failure Handling
Diagnosis should fail soft.
If probes cannot reach a strong conclusion, the output should still be structured:
GateDiagnosis
primary_failure_class = unknown
recommended_probes = [...]
recommended_interventions = [collect_more_evidence]
confidence = low
That is better than forcing a false diagnosis.
Why This Matters
This execution flow is what keeps whetstone_RSA from becoming just a model zoo.
The library’s value is not only that it can run tiny models. It is that it can help decide:
- whether a gate is worth scaling
- whether a gate should be restructured
- whether the problem should leave ML entirely