Initial whetstone_RSA architecture and C++ scaffold
This commit is contained in:
290
docs/probe_execution_flow.md
Normal file
290
docs/probe_execution_flow.md
Normal file
@@ -0,0 +1,290 @@
|
||||
# Probe Execution Flow
|
||||
|
||||
## Purpose
|
||||
|
||||
This document describes how `whetstone_RSA` should execute gate diagnosis in
|
||||
practice.
|
||||
|
||||
The taxonomy and diagnosis API define what the library needs to represent.
|
||||
This document defines how the pieces fit together during a diagnosis run.
|
||||
|
||||
## High-Level Flow
|
||||
|
||||
```text
|
||||
GateDefinition
|
||||
+
|
||||
GateEvidence
|
||||
+
|
||||
PolicyContext
|
||||
↓
|
||||
Probe selection
|
||||
↓
|
||||
Applicable probes run
|
||||
↓
|
||||
Probe results normalized
|
||||
↓
|
||||
Failure-class synthesis
|
||||
↓
|
||||
Intervention ranking
|
||||
↓
|
||||
Policy recommendation
|
||||
↓
|
||||
GateDiagnosis
|
||||
```
|
||||
|
||||
## Inputs
|
||||
|
||||
### 1. `GateDefinition`
|
||||
|
||||
Static metadata about the gate:
|
||||
|
||||
- labels
|
||||
- slots
|
||||
- constraints
|
||||
- risk tier
|
||||
- supports abstain
|
||||
- deterministic baseline availability
|
||||
- label stability
|
||||
- candidate factorizations
|
||||
|
||||
### 2. `GateEvidence`
|
||||
|
||||
Observed evidence for one gate and one model/policy configuration:
|
||||
|
||||
- raw accuracy
|
||||
- accepted accuracy
|
||||
- abstain rate
|
||||
- retry recovery
|
||||
- silent error rate
|
||||
- calibration
|
||||
- confusion summary
|
||||
- latency
|
||||
- cost
|
||||
|
||||
### 3. `PolicyContext`
|
||||
|
||||
Deployment context that affects recommendations:
|
||||
|
||||
- risk budget
|
||||
- acceptable latency
|
||||
- acceptable compute cost
|
||||
- available escalation targets
|
||||
- whether deterministic checks exist downstream
|
||||
|
||||
## Step 1: Validate Evidence
|
||||
|
||||
Before diagnosis starts, the library should verify the evidence is usable.
|
||||
|
||||
Checks:
|
||||
|
||||
- gate id matches between definition and evidence
|
||||
- labels in confusion summary match the schema
|
||||
- metrics are internally consistent
|
||||
- risk tier and policy context are present
|
||||
|
||||
If evidence is incomplete, diagnosis should still proceed when possible, but the
|
||||
missing fields should reduce confidence in the result.
|
||||
|
||||
## Step 2: Select Applicable Probes
|
||||
|
||||
Not every probe applies to every gate.
|
||||
|
||||
Examples:
|
||||
|
||||
- `deterministic_rule_probe` applies only when structured features or a deterministic
|
||||
baseline are available
|
||||
- `factorization_probe` applies when labels suggest composition or a candidate
|
||||
factorization is registered
|
||||
- `schema_stability_probe` applies when labels may drift with policy or capability
|
||||
- `confidence_threshold_probe` applies only if confidence-bearing outputs exist
|
||||
|
||||
The library should ask each probe:
|
||||
|
||||
```text
|
||||
supports(gate_definition, evidence) -> bool
|
||||
```
|
||||
|
||||
and only run probes that opt in.
|
||||
|
||||
## Step 3: Run Probes
|
||||
|
||||
Each applicable probe returns:
|
||||
|
||||
- probe-specific signals
|
||||
- recommendation hints
|
||||
- a confidence score
|
||||
|
||||
Probe outputs should be normalized into a standard envelope before synthesis.
|
||||
|
||||
Example normalized structure:
|
||||
|
||||
```text
|
||||
ProbeResult
|
||||
probe_id
|
||||
status
|
||||
signals
|
||||
recommendation_hints
|
||||
confidence
|
||||
```
|
||||
|
||||
## Step 4: Aggregate Signals
|
||||
|
||||
The library should not map one probe directly to one final diagnosis.
|
||||
|
||||
Instead it should synthesize across probe outputs.
|
||||
|
||||
Examples:
|
||||
|
||||
- strong `factorization_probe` + compositional confusion pattern
|
||||
-> `factorizable_gate`
|
||||
- strong `schema_stability_probe` + drifting label semantics
|
||||
-> `non_stationary_gate`
|
||||
- strong `confidence_threshold_probe` + weak raw accuracy but high accepted precision
|
||||
-> `guardrail_limited_gate`
|
||||
- strong `deterministic_rule_probe`
|
||||
-> `deterministic_disguised_as_ml_gate`
|
||||
|
||||
This step should produce:
|
||||
|
||||
- ranked failure classes
|
||||
- supporting signals
|
||||
- unresolved ambiguities
|
||||
|
||||
## Step 5: Rank Interventions
|
||||
|
||||
Once failure classes are ranked, the library should rank interventions.
|
||||
|
||||
Interventions should come from a controlled vocabulary:
|
||||
|
||||
- `keep_current_gate`
|
||||
- `deploy_with_guardrails`
|
||||
- `raise_confidence_threshold`
|
||||
- `add_retry_with_context`
|
||||
- `add_deterministic_prepass`
|
||||
- `replace_with_deterministic_logic`
|
||||
- `factor_into_subgates`
|
||||
- `switch_to_hierarchical_gate`
|
||||
- `revise_label_schema`
|
||||
- `enrich_input_context`
|
||||
- `evaluate_larger_model_tier`
|
||||
|
||||
The order matters.
|
||||
|
||||
The library should prefer lower-cost, structurally-corrective interventions before
|
||||
defaulting to bigger models.
|
||||
|
||||
## Step 6: Emit Policy Recommendation
|
||||
|
||||
The final diagnosis should include a recommended deployment posture.
|
||||
|
||||
Possible deployment modes:
|
||||
|
||||
- `auto_accept`
|
||||
- `guarded_accept`
|
||||
- `abstain_first`
|
||||
- `research_only`
|
||||
|
||||
This recommendation should consider:
|
||||
|
||||
- risk tier
|
||||
- silent error estimate
|
||||
- accepted precision
|
||||
- latency budget
|
||||
- escalation availability
|
||||
|
||||
## Example Flow Patterns
|
||||
|
||||
### Pattern A: Good Gate With Policy Gaps
|
||||
|
||||
Observed:
|
||||
|
||||
- moderate raw accuracy
|
||||
- strong accepted precision under thresholding
|
||||
- downstream deterministic checks exist
|
||||
|
||||
Likely result:
|
||||
|
||||
- primary failure class: `guardrail_limited_gate`
|
||||
- intervention: `deploy_with_guardrails`
|
||||
|
||||
### Pattern B: Flat Multiclass Gate
|
||||
|
||||
Observed:
|
||||
|
||||
- low raw accuracy
|
||||
- confusion concentrates between composite labels
|
||||
- factorization metadata exists
|
||||
|
||||
Likely result:
|
||||
|
||||
- primary failure class: `factorizable_gate`
|
||||
- intervention: `factor_into_subgates`
|
||||
|
||||
### Pattern C: Formula-Like Gate
|
||||
|
||||
Observed:
|
||||
|
||||
- high score under structural encoding
|
||||
- symbolic baseline exists
|
||||
- class meaning is crisp
|
||||
|
||||
Likely result:
|
||||
|
||||
- primary failure class: `deterministic_disguised_as_ml_gate`
|
||||
- intervention: `replace_with_deterministic_logic`
|
||||
|
||||
### Pattern D: Drifting Capability Boundary
|
||||
|
||||
Observed:
|
||||
|
||||
- disagreements cluster between adjacent routing tiers
|
||||
- label meanings depend on current tooling maturity
|
||||
|
||||
Likely result:
|
||||
|
||||
- primary failure class: `non_stationary_gate`
|
||||
- intervention: `revise_label_schema` or separate ontology from policy
|
||||
|
||||
## Confidence Model
|
||||
|
||||
The diagnosis should carry its own confidence.
|
||||
|
||||
Diagnosis confidence should depend on:
|
||||
|
||||
- probe coverage
|
||||
- evidence completeness
|
||||
- signal agreement across probes
|
||||
- stability of the ranked recommendation
|
||||
|
||||
Low-confidence diagnosis is still useful if it says:
|
||||
|
||||
- evidence insufficient
|
||||
- collect richer context
|
||||
- run a missing probe
|
||||
|
||||
## Failure Handling
|
||||
|
||||
Diagnosis should fail soft.
|
||||
|
||||
If probes cannot reach a strong conclusion, the output should still be structured:
|
||||
|
||||
```text
|
||||
GateDiagnosis
|
||||
primary_failure_class = unknown
|
||||
recommended_probes = [...]
|
||||
recommended_interventions = [collect_more_evidence]
|
||||
confidence = low
|
||||
```
|
||||
|
||||
That is better than forcing a false diagnosis.
|
||||
|
||||
## Why This Matters
|
||||
|
||||
This execution flow is what keeps `whetstone_RSA` from becoming just a model zoo.
|
||||
|
||||
The library’s value is not only that it can run tiny models. It is that it can
|
||||
help decide:
|
||||
|
||||
- whether a gate is worth scaling
|
||||
- whether a gate should be restructured
|
||||
- whether the problem should leave ML entirely
|
||||
Reference in New Issue
Block a user