5.0 KiB
Runtime Contract
Purpose
This document defines the minimal runtime contract for whetstone_RSA.
The project already has:
- gate schemas
- evaluation policy
- failure taxonomy
- diagnosis API sketch
- probe execution flow
The missing piece is the execution boundary:
- what the runtime receives
- what it returns
- what extension points exist
This document defines that contract without binding the project to a specific implementation language.
Runtime Roles
The runtime has three primary roles:
1. Decision Runtime
Runs a bounded gate against an input request and returns a structured decision.
2. Diagnosis Runtime
Runs probes against gate evidence and returns a structured diagnosis.
3. Policy Runtime
Combines gate output and diagnosis output into a deployment posture:
- accept
- abstain
- retry
- escalate
Backend selection is a related but distinct concern and should remain adjacent to
the RSA runtime rather than defining the RSA core. See docs/backend_selection.md.
Core Runtime Calls
decide
Purpose:
Run a bounded decision gate.
Conceptual signature:
decide(decision_request) -> decision_result
Input:
DecisionRequest
gate_id
input
context
policy_context
Output:
DecisionResult
gate_id
decision_id
slots
confidence
abstain
escalation_target
trace
diagnose
Purpose:
Run gate diagnosis based on metadata and evidence.
Conceptual signature:
diagnose(diagnosis_request) -> gate_diagnosis
Input:
DiagnosisRequest
gate_definition
evidence
policy_context
enabled_probes
Output:
GateDiagnosis
gate_id
primary_failure_class
secondary_failure_classes
supporting_signals
recommended_probes
recommended_interventions
recommended_policy
confidence
recommend_policy
Purpose:
Convert a gate decision and a diagnosis into an action posture.
Conceptual signature:
recommend_policy(decision_result, gate_diagnosis, policy_context) -> policy_action
Output:
PolicyAction
mode
threshold_applied
retry_strategy
escalation_target
rationale
Runtime Responsibilities
The runtime should own:
- schema validation
- gate lookup
- model-tier lookup
- probe lookup
- evidence normalization
- policy recommendation
The runtime should not own:
- domain-specific business logic after execution
- arbitrary open-ended generation
- hidden manual policy that bypasses the declared schema
Extension Points
The runtime should expose three registries:
Gate Registry
Maps gate_id to:
- gate definition
- supported model tiers
- deterministic baseline if any
- candidate factorizations
Probe Registry
Maps probe_id to:
- applicability rules
- execution logic
- output normalizer
Policy Registry
Maps deployment mode or risk policy to:
- threshold rules
- retry rules
- escalation rules
Minimal Validation Rules
For decide
- gate must exist
- decision schema must validate
- model tier must be available if the gate is learned
- inputs must satisfy required structure
For diagnose
- gate definition must validate
- evidence must validate
- enabled probes must exist
- probe applicability must be checked before execution
For recommend_policy
- decision result and diagnosis must refer to the same gate
- risk tier must be available
- policy mode must be valid for the gate
Trace Requirements
The runtime should preserve structured trace output.
Decision Trace
Should include:
- selected gate
- selected model tier
- confidence
- thresholding outcome
- abstain or escalation outcome
Diagnosis Trace
Should include:
- probes run
- probe outputs
- synthesized failure class ranking
- intervention ranking
Trace should be inspectable by humans and consumable by tools.
Recommended Runtime States
For implementation planning, treat runtime work as these internal stages:
request_received
schema_validated
gate_loaded
decision_executed
evidence_loaded
probes_executed
diagnosis_synthesized
policy_computed
response_emitted
This is useful for debugging and logging even before the runtime is complex.
Failure Handling
The runtime should fail in a structured way.
Example error kinds:
unknown_gateinvalid_schemamissing_model_tierinvalid_evidenceunknown_probeinapplicable_probediagnosis_insufficient_evidence
Errors should be explicit objects, not plain strings.
Implementation Guidance
The first implementation should stay minimal:
- one gate registry abstraction
- one probe registry abstraction
- one diagnosis synthesis path
- one JSON-serializable result format
Do not overbuild orchestration first.
The value of the runtime is that it makes bounded gate execution and diagnosis uniform across applications.
Relationship To Case Studies
Case studies should populate the runtime through data:
- gate definitions
- evidence bundles
- diagnosis outputs
They should not force case-specific branches into the runtime contract.