Files
whetstone_RSA/docs/probe_execution_flow.md
2026-03-31 22:50:40 -06:00

6.2 KiB
Raw Permalink Blame History

Probe Execution Flow

Purpose

This document describes how whetstone_RSA should execute gate diagnosis in practice.

The taxonomy and diagnosis API define what the library needs to represent. This document defines how the pieces fit together during a diagnosis run.

High-Level Flow

GateDefinition
    +
GateEvidence
    +
PolicyContext
    ↓
Probe selection
    ↓
Applicable probes run
    ↓
Probe results normalized
    ↓
Failure-class synthesis
    ↓
Intervention ranking
    ↓
Policy recommendation
    ↓
GateDiagnosis

Inputs

1. GateDefinition

Static metadata about the gate:

  • labels
  • slots
  • constraints
  • risk tier
  • supports abstain
  • deterministic baseline availability
  • label stability
  • candidate factorizations

2. GateEvidence

Observed evidence for one gate and one model/policy configuration:

  • raw accuracy
  • accepted accuracy
  • abstain rate
  • retry recovery
  • silent error rate
  • calibration
  • confusion summary
  • latency
  • cost

3. PolicyContext

Deployment context that affects recommendations:

  • risk budget
  • acceptable latency
  • acceptable compute cost
  • available escalation targets
  • whether deterministic checks exist downstream

Step 1: Validate Evidence

Before diagnosis starts, the library should verify the evidence is usable.

Checks:

  • gate id matches between definition and evidence
  • labels in confusion summary match the schema
  • metrics are internally consistent
  • risk tier and policy context are present

If evidence is incomplete, diagnosis should still proceed when possible, but the missing fields should reduce confidence in the result.

Step 2: Select Applicable Probes

Not every probe applies to every gate.

Examples:

  • deterministic_rule_probe applies only when structured features or a deterministic baseline are available
  • factorization_probe applies when labels suggest composition or a candidate factorization is registered
  • schema_stability_probe applies when labels may drift with policy or capability
  • confidence_threshold_probe applies only if confidence-bearing outputs exist

The library should ask each probe:

supports(gate_definition, evidence) -> bool

and only run probes that opt in.

Step 3: Run Probes

Each applicable probe returns:

  • probe-specific signals
  • recommendation hints
  • a confidence score

Probe outputs should be normalized into a standard envelope before synthesis.

Example normalized structure:

ProbeResult
  probe_id
  status
  signals
  recommendation_hints
  confidence

Step 4: Aggregate Signals

The library should not map one probe directly to one final diagnosis.

Instead it should synthesize across probe outputs.

Examples:

  • strong factorization_probe + compositional confusion pattern -> factorizable_gate
  • strong schema_stability_probe + drifting label semantics -> non_stationary_gate
  • strong confidence_threshold_probe + weak raw accuracy but high accepted precision -> guardrail_limited_gate
  • strong deterministic_rule_probe -> deterministic_disguised_as_ml_gate

This step should produce:

  • ranked failure classes
  • supporting signals
  • unresolved ambiguities

Step 5: Rank Interventions

Once failure classes are ranked, the library should rank interventions.

Interventions should come from a controlled vocabulary:

  • keep_current_gate
  • deploy_with_guardrails
  • raise_confidence_threshold
  • add_retry_with_context
  • add_deterministic_prepass
  • replace_with_deterministic_logic
  • factor_into_subgates
  • switch_to_hierarchical_gate
  • revise_label_schema
  • enrich_input_context
  • evaluate_larger_model_tier

The order matters.

The library should prefer lower-cost, structurally-corrective interventions before defaulting to bigger models.

Step 6: Emit Policy Recommendation

The final diagnosis should include a recommended deployment posture.

Possible deployment modes:

  • auto_accept
  • guarded_accept
  • abstain_first
  • research_only

This recommendation should consider:

  • risk tier
  • silent error estimate
  • accepted precision
  • latency budget
  • escalation availability

Example Flow Patterns

Pattern A: Good Gate With Policy Gaps

Observed:

  • moderate raw accuracy
  • strong accepted precision under thresholding
  • downstream deterministic checks exist

Likely result:

  • primary failure class: guardrail_limited_gate
  • intervention: deploy_with_guardrails

Pattern B: Flat Multiclass Gate

Observed:

  • low raw accuracy
  • confusion concentrates between composite labels
  • factorization metadata exists

Likely result:

  • primary failure class: factorizable_gate
  • intervention: factor_into_subgates

Pattern C: Formula-Like Gate

Observed:

  • high score under structural encoding
  • symbolic baseline exists
  • class meaning is crisp

Likely result:

  • primary failure class: deterministic_disguised_as_ml_gate
  • intervention: replace_with_deterministic_logic

Pattern D: Drifting Capability Boundary

Observed:

  • disagreements cluster between adjacent routing tiers
  • label meanings depend on current tooling maturity

Likely result:

  • primary failure class: non_stationary_gate
  • intervention: revise_label_schema or separate ontology from policy

Confidence Model

The diagnosis should carry its own confidence.

Diagnosis confidence should depend on:

  • probe coverage
  • evidence completeness
  • signal agreement across probes
  • stability of the ranked recommendation

Low-confidence diagnosis is still useful if it says:

  • evidence insufficient
  • collect richer context
  • run a missing probe

Failure Handling

Diagnosis should fail soft.

If probes cannot reach a strong conclusion, the output should still be structured:

GateDiagnosis
  primary_failure_class = unknown
  recommended_probes = [...]
  recommended_interventions = [collect_more_evidence]
  confidence = low

That is better than forcing a false diagnosis.

Why This Matters

This execution flow is what keeps whetstone_RSA from becoming just a model zoo.

The librarys value is not only that it can run tiny models. It is that it can help decide:

  • whether a gate is worth scaling
  • whether a gate should be restructured
  • whether the problem should leave ML entirely