WIP: save progress 2026-04-12
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
299
ROADMAP.md
299
ROADMAP.md
@@ -6,12 +6,115 @@
|
||||
- Define the bounded gate schema contract.
|
||||
- Document the WhetstoneDSL case study and current specialist fleet.
|
||||
|
||||
### Phase 0A: AST-First Realignment
|
||||
|
||||
Before `whetstone_RSA` grows much further as a hand-authored C++ library, realign
|
||||
the project so its semantic source of truth is Whetstone-authored and C++ becomes
|
||||
a projection target.
|
||||
|
||||
This phase exists to prevent the current bootstrap scaffold from hardening into the
|
||||
wrong long-term architecture.
|
||||
|
||||
#### Intent
|
||||
|
||||
- make RSA semantically AST-first
|
||||
- keep C++ as runtime projection and execution layer
|
||||
- preserve compatibility with the current scaffold while changing the source of truth
|
||||
|
||||
#### Architectural rule
|
||||
|
||||
The long-term source of truth for RSA should live in Whetstone semantic forms, not
|
||||
in hand-authored C++ structs.
|
||||
|
||||
The C++ code in this repository should be treated as:
|
||||
|
||||
- runtime implementation
|
||||
- projected contract types
|
||||
- backend execution surface
|
||||
- diagnosis and policy engine
|
||||
|
||||
not the canonical definition of RSA semantics.
|
||||
|
||||
#### 0A.1 Semantic entities to define first
|
||||
|
||||
Define Whetstone-authored semantic objects for at least:
|
||||
|
||||
- `GateDefinition`
|
||||
- `DecisionRequest`
|
||||
- `DecisionResult`
|
||||
- `StructuredDecisionResult`
|
||||
- `GateSuitability`
|
||||
- `GateDiagnosis`
|
||||
- `ProbeDefinition`
|
||||
- `BackendSuitability`
|
||||
- `GatePackage`
|
||||
|
||||
These should carry the contract metadata already identified in the docs:
|
||||
|
||||
- input structure and modality
|
||||
- output topology and constraint
|
||||
- policy stability
|
||||
- gate shape
|
||||
- confidence, abstain, and escalation policy
|
||||
- trace fields
|
||||
- backend family and packaging metadata
|
||||
|
||||
#### 0A.2 Projection boundary
|
||||
|
||||
Define a deliberate projection boundary from Whetstone semantic forms into:
|
||||
|
||||
- C++ runtime contract structs
|
||||
- schema artifacts
|
||||
- benchmark configuration artifacts
|
||||
- packaged runtime manifests
|
||||
|
||||
The important architectural constraint is:
|
||||
|
||||
- semantics are authored once
|
||||
- runtime targets are projected from that source
|
||||
|
||||
#### 0A.3 Migration rule for the current scaffold
|
||||
|
||||
Do not continue expanding hand-authored C++ contract definitions unless they are
|
||||
clearly temporary runtime placeholders.
|
||||
|
||||
When a contract type changes, the preferred direction should become:
|
||||
|
||||
- update semantic source
|
||||
- regenerate or realign projected C++ representation
|
||||
|
||||
rather than:
|
||||
|
||||
- evolve C++ first
|
||||
- treat docs and later AST work as catch-up
|
||||
|
||||
#### 0A.4 Immediate AST-first deliverables
|
||||
|
||||
- define the first RSA semantic model in Whetstone terms
|
||||
- map current C++ runtime types to those semantic entities
|
||||
- identify what current fields are semantic contract versus runtime-only detail
|
||||
- define the projection rules into C++
|
||||
- update roadmap phases so future implementation assumes semantic-first authoring
|
||||
|
||||
#### 0A.5 Scope constraint
|
||||
|
||||
This does not require immediate deep modification of `whetstone_DSL` runtime code.
|
||||
|
||||
It does require that future RSA design work stop assuming C++ is the authoring
|
||||
home of the system.
|
||||
|
||||
The immediate goal is architectural realignment, not premature full conversion.
|
||||
|
||||
## Phase 1: Gate Inventory
|
||||
|
||||
- Enumerate candidate gates from WhetstoneDSL.
|
||||
- Record output vocabulary, input shape, and downstream deterministic executor.
|
||||
- Split gates into deterministic, specialist, and escalation-first candidates.
|
||||
|
||||
Primary case-study inventory artifact:
|
||||
|
||||
- `docs/case_studies/whetstone_dsl_gate_inventory.md`
|
||||
|
||||
## Phase 2: Entropy Baseline
|
||||
|
||||
- Define a first-pass gate entropy score.
|
||||
@@ -31,6 +134,194 @@
|
||||
- Implement abstain and escalation policies.
|
||||
- Support pluggable inference backends.
|
||||
|
||||
### Phase 4A: Post-Suitability RSA Runtime Plan
|
||||
|
||||
This phase begins only after a decision surface has already been judged to belong
|
||||
in the `rsa` layer rather than `deterministic`, `slm`, `llm`, or `human`.
|
||||
|
||||
The focus here is not broad suitability triage. The focus is:
|
||||
|
||||
- how an RSA gate is represented
|
||||
- how its shape is chosen
|
||||
- how its backend is selected
|
||||
- how it is packaged for efficient runtime use
|
||||
|
||||
#### Design constraints
|
||||
|
||||
- Keep the library focused on `whetstone_RSA`, not direct `whetstone_DSL` changes.
|
||||
- Assume RSA semantic contracts are authored AST-first and projected into runtime
|
||||
targets rather than defined primarily in handwritten C++.
|
||||
- Support both bounded single-choice classification and bounded structured outputs
|
||||
with slots from the beginning.
|
||||
- Include `rl_policy` and bandit-style policies as contract placeholders, but do
|
||||
not make them a primary implementation target until classical ML and transformer
|
||||
backends are working and benchmarked.
|
||||
- Treat hybrid inputs as first-class. A gate may consume text, symbolic state, and
|
||||
numeric features together.
|
||||
- Use case studies only to surface reusable abstractions. Do not let one case study
|
||||
become the semantic definition of the library.
|
||||
|
||||
#### 4A.1 Contract and packet model
|
||||
|
||||
- Define a canonical `GateDefinition` that supports:
|
||||
- bounded enum decisions
|
||||
- bounded structured slot-filling decisions
|
||||
- constraints
|
||||
- confidence policy
|
||||
- abstain policy
|
||||
- escalation policy
|
||||
- trace metadata
|
||||
- Define a canonical `DecisionRequest` packet that supports:
|
||||
- textual fields
|
||||
- symbolic fields
|
||||
- numeric features
|
||||
- state and provenance metadata
|
||||
- Define `DecisionResult` and `StructuredDecisionResult` so programs can consume
|
||||
either plain label outputs or bounded slot outputs through argument-oriented
|
||||
runtime calls rather than project-specific ad hoc interfaces.
|
||||
|
||||
#### 4A.2 Gate-shape selection
|
||||
|
||||
The library should not assume that all bounded problems should be represented as
|
||||
one flat multiclass classifier.
|
||||
|
||||
For each RSA candidate gate, support shape analysis across:
|
||||
|
||||
- `binary`
|
||||
- `multiclass`
|
||||
- `factorized`
|
||||
- `hierarchical`
|
||||
|
||||
Initial implementation goals:
|
||||
|
||||
- define metadata that describes candidate gate shapes
|
||||
- define diagnostics that detect when a flat label space hides compositional
|
||||
structure
|
||||
- support shape-specific constraints so invalid state combinations are prevented
|
||||
by contract rather than patched later
|
||||
|
||||
#### 4A.3 Backend-selection layer
|
||||
|
||||
Backend selection remains separate from RSA identity.
|
||||
|
||||
The library should support a backend contract for at least:
|
||||
|
||||
- `linear_model`
|
||||
- `gradient_boosting`
|
||||
- `random_forest`
|
||||
- `bayesian`
|
||||
- `transformer`
|
||||
- `hybrid`
|
||||
- `rl_policy` placeholder
|
||||
|
||||
Initial implementation goals:
|
||||
|
||||
- define a common backend interface for train, infer, calibrate, and trace
|
||||
- allow external benchmark or AutoML-style workflows to produce backend
|
||||
recommendations without forcing one internal training stack
|
||||
- keep backend artifacts swappable as long as the gate contract is preserved
|
||||
|
||||
#### 4A.4 Benchmark harness
|
||||
|
||||
Build a reusable harness that compares both gate shape and backend family for the
|
||||
same bounded decision contract.
|
||||
|
||||
For each gate, the harness should compare:
|
||||
|
||||
- flat multiclass against factorized or hierarchical alternatives
|
||||
- classical ML backends against tiny transformer backends
|
||||
- deterministic baselines where any part of the gate can be rule-complete
|
||||
|
||||
Required outputs:
|
||||
|
||||
- accepted-decision accuracy
|
||||
- calibration quality
|
||||
- abstain quality
|
||||
- invalid-output rate
|
||||
- downstream deterministic validation success
|
||||
- latency
|
||||
- memory footprint
|
||||
- training and retraining cost
|
||||
|
||||
#### 4A.5 Runtime packaging
|
||||
|
||||
AutoML or scripted comparison should be offline or occasional, not on the hot path.
|
||||
|
||||
The runtime should load a selected gate package that includes:
|
||||
|
||||
- gate contract
|
||||
- chosen gate shape
|
||||
- chosen backend family
|
||||
- trained artifact reference
|
||||
- confidence thresholds
|
||||
- abstain and escalation policy
|
||||
- trace schema and version metadata
|
||||
|
||||
The steady-state runtime model is:
|
||||
|
||||
- benchmark and select offline
|
||||
- package once
|
||||
- load the chosen gate into memory
|
||||
- serve repeated decisions efficiently
|
||||
- re-run selection only when schema, data, policy, or hardware conditions change
|
||||
|
||||
#### 4A.6 Invocation surface
|
||||
|
||||
The library should expose explicit runtime calls rather than force consumers to
|
||||
rebuild invocation conventions for each gate.
|
||||
|
||||
Initial target surface:
|
||||
|
||||
- `decide_class(request, gate_definition)`
|
||||
- `decide_structured(request, gate_definition)`
|
||||
- `diagnose_gate(gate_definition, evidence, benchmark_results)`
|
||||
- `recommend_backend(gate_definition, benchmark_results)`
|
||||
- `recommend_gate_shape(gate_definition, benchmark_results)`
|
||||
|
||||
The exact function names may change, but the core idea should remain: consumers
|
||||
pass structured arguments into a reusable RSA runtime surface.
|
||||
|
||||
#### 4A.7 Diagnosis loop
|
||||
|
||||
When a gate underperforms, the first response should not be automatic model-size
|
||||
escalation.
|
||||
|
||||
The runtime and benchmark surfaces should instead help answer:
|
||||
|
||||
- should this gate stay flat or be factorized?
|
||||
- is part of the task actually deterministic?
|
||||
- is the label space unstable?
|
||||
- is the backend family mismatched to the input packet?
|
||||
- should the gate abstain earlier?
|
||||
|
||||
#### 4A.8 First case-study role
|
||||
|
||||
Use the memory-strategy decision only as the first end-to-end exercise for the
|
||||
abstract mechanism.
|
||||
|
||||
The case study should help validate:
|
||||
|
||||
- mixed symbolic/text/numeric packet design
|
||||
- shape comparison between flat and decomposed gates
|
||||
- backend comparison between classical ML and tiny transformers
|
||||
- packaging of both enum-style and slot-filling outputs
|
||||
|
||||
It should not be allowed to overfit the library to one memory-management ontology.
|
||||
|
||||
#### 4A.9 Immediate implementation order
|
||||
|
||||
- define the AST-first semantic source for RSA contracts
|
||||
- define projection rules from semantic entities into C++ runtime types
|
||||
- extend schema types so both enum and structured-slot gates are first-class
|
||||
- define canonical request/result packet types
|
||||
- define gate-shape metadata and recommendation types
|
||||
- define backend-family contracts, including placeholder RL/bandit support
|
||||
- build the benchmark report format
|
||||
- implement classical ML adapters first
|
||||
- implement transformer adapter second
|
||||
- add packaging and runtime loading
|
||||
- add diagnosis and recommendation APIs on top of benchmark evidence
|
||||
|
||||
## Phase 5: Whetstone Integration
|
||||
|
||||
- Expose WhetstoneDSL gates through the library.
|
||||
@@ -46,6 +337,12 @@
|
||||
## Exit Criteria For Version 0
|
||||
|
||||
- At least one reusable gate schema format is stable.
|
||||
- Both enum-style and bounded structured-slot gates are supported by the contract.
|
||||
- At least three WhetstoneDSL gates are benchmarked across size tiers.
|
||||
- At least one gate has been compared across multiple gate shapes, not just
|
||||
multiple model sizes.
|
||||
- At least one gate has been compared across classical ML and transformer
|
||||
backends under the same contract.
|
||||
- A baseline entropy-to-tier recommendation table exists.
|
||||
- One real application can call the library instead of hard-coding model selection.
|
||||
- One real application can call the library instead of hard-coding model
|
||||
selection or backend choice.
|
||||
|
||||
Reference in New Issue
Block a user