# Roadmap ## Phase 0: Definition - Write the project framing and library boundaries. - Define the bounded gate schema contract. - Document the WhetstoneDSL case study and current specialist fleet. ### Phase 0A: AST-First Realignment Before `whetstone_RSA` grows much further as a hand-authored C++ library, realign the project so its semantic source of truth is Whetstone-authored and C++ becomes a projection target. This phase exists to prevent the current bootstrap scaffold from hardening into the wrong long-term architecture. #### Intent - make RSA semantically AST-first - keep C++ as runtime projection and execution layer - preserve compatibility with the current scaffold while changing the source of truth #### Architectural rule The long-term source of truth for RSA should live in Whetstone semantic forms, not in hand-authored C++ structs. The C++ code in this repository should be treated as: - runtime implementation - projected contract types - backend execution surface - diagnosis and policy engine not the canonical definition of RSA semantics. #### 0A.1 Semantic entities to define first Define Whetstone-authored semantic objects for at least: - `GateDefinition` - `DecisionRequest` - `DecisionResult` - `StructuredDecisionResult` - `GateSuitability` - `GateDiagnosis` - `ProbeDefinition` - `BackendSuitability` - `GatePackage` These should carry the contract metadata already identified in the docs: - input structure and modality - output topology and constraint - policy stability - gate shape - confidence, abstain, and escalation policy - trace fields - backend family and packaging metadata #### 0A.2 Projection boundary Define a deliberate projection boundary from Whetstone semantic forms into: - C++ runtime contract structs - schema artifacts - benchmark configuration artifacts - packaged runtime manifests The important architectural constraint is: - semantics are authored once - runtime targets are projected from that source #### 0A.3 Migration rule for the current scaffold Do not continue expanding hand-authored C++ contract definitions unless they are clearly temporary runtime placeholders. When a contract type changes, the preferred direction should become: - update semantic source - regenerate or realign projected C++ representation rather than: - evolve C++ first - treat docs and later AST work as catch-up #### 0A.4 Immediate AST-first deliverables - define the first RSA semantic model in Whetstone terms - map current C++ runtime types to those semantic entities - identify what current fields are semantic contract versus runtime-only detail - define the projection rules into C++ - update roadmap phases so future implementation assumes semantic-first authoring #### 0A.5 Scope constraint This does not require immediate deep modification of `whetstone_DSL` runtime code. It does require that future RSA design work stop assuming C++ is the authoring home of the system. The immediate goal is architectural realignment, not premature full conversion. ## Phase 1: Gate Inventory - Enumerate candidate gates from WhetstoneDSL. - Record output vocabulary, input shape, and downstream deterministic executor. - Split gates into deterministic, specialist, and escalation-first candidates. Primary case-study inventory artifact: - `docs/case_studies/whetstone_dsl_gate_inventory.md` ## Phase 2: Entropy Baseline - Define a first-pass gate entropy score. - Score the initial WhetstoneDSL gate set. - Identify which gates are plausible at the ~800KB tier and which are not. ## Phase 3: Size Sweep - Run fixed-data comparisons across multiple model sizes. - Measure accuracy, calibration, confusion concentration, and latency. - Produce the first entropy-to-tier baseline table. ## Phase 4: Core Library - Implement gate schema loading and validation. - Implement the decision request/response runtime. - Implement abstain and escalation policies. - Support pluggable inference backends. ### Phase 4A: Post-Suitability RSA Runtime Plan This phase begins only after a decision surface has already been judged to belong in the `rsa` layer rather than `deterministic`, `slm`, `llm`, or `human`. The focus here is not broad suitability triage. The focus is: - how an RSA gate is represented - how its shape is chosen - how its backend is selected - how it is packaged for efficient runtime use #### Design constraints - Keep the library focused on `whetstone_RSA`, not direct `whetstone_DSL` changes. - Assume RSA semantic contracts are authored AST-first and projected into runtime targets rather than defined primarily in handwritten C++. - Support both bounded single-choice classification and bounded structured outputs with slots from the beginning. - Include `rl_policy` and bandit-style policies as contract placeholders, but do not make them a primary implementation target until classical ML and transformer backends are working and benchmarked. - Treat hybrid inputs as first-class. A gate may consume text, symbolic state, and numeric features together. - Use case studies only to surface reusable abstractions. Do not let one case study become the semantic definition of the library. #### 4A.1 Contract and packet model - Define a canonical `GateDefinition` that supports: - bounded enum decisions - bounded structured slot-filling decisions - constraints - confidence policy - abstain policy - escalation policy - trace metadata - Define a canonical `DecisionRequest` packet that supports: - textual fields - symbolic fields - numeric features - state and provenance metadata - Define `DecisionResult` and `StructuredDecisionResult` so programs can consume either plain label outputs or bounded slot outputs through argument-oriented runtime calls rather than project-specific ad hoc interfaces. #### 4A.2 Gate-shape selection The library should not assume that all bounded problems should be represented as one flat multiclass classifier. For each RSA candidate gate, support shape analysis across: - `binary` - `multiclass` - `factorized` - `hierarchical` Initial implementation goals: - define metadata that describes candidate gate shapes - define diagnostics that detect when a flat label space hides compositional structure - support shape-specific constraints so invalid state combinations are prevented by contract rather than patched later #### 4A.3 Backend-selection layer Backend selection remains separate from RSA identity. The library should support a backend contract for at least: - `linear_model` - `gradient_boosting` - `random_forest` - `bayesian` - `transformer` - `hybrid` - `rl_policy` placeholder Initial implementation goals: - define a common backend interface for train, infer, calibrate, and trace - allow external benchmark or AutoML-style workflows to produce backend recommendations without forcing one internal training stack - keep backend artifacts swappable as long as the gate contract is preserved #### 4A.4 Benchmark harness Build a reusable harness that compares both gate shape and backend family for the same bounded decision contract. For each gate, the harness should compare: - flat multiclass against factorized or hierarchical alternatives - classical ML backends against tiny transformer backends - deterministic baselines where any part of the gate can be rule-complete Required outputs: - accepted-decision accuracy - calibration quality - abstain quality - invalid-output rate - downstream deterministic validation success - latency - memory footprint - training and retraining cost #### 4A.5 Runtime packaging AutoML or scripted comparison should be offline or occasional, not on the hot path. The runtime should load a selected gate package that includes: - gate contract - chosen gate shape - chosen backend family - trained artifact reference - confidence thresholds - abstain and escalation policy - trace schema and version metadata The steady-state runtime model is: - benchmark and select offline - package once - load the chosen gate into memory - serve repeated decisions efficiently - re-run selection only when schema, data, policy, or hardware conditions change #### 4A.6 Invocation surface The library should expose explicit runtime calls rather than force consumers to rebuild invocation conventions for each gate. Initial target surface: - `decide_class(request, gate_definition)` - `decide_structured(request, gate_definition)` - `diagnose_gate(gate_definition, evidence, benchmark_results)` - `recommend_backend(gate_definition, benchmark_results)` - `recommend_gate_shape(gate_definition, benchmark_results)` The exact function names may change, but the core idea should remain: consumers pass structured arguments into a reusable RSA runtime surface. #### 4A.7 Diagnosis loop When a gate underperforms, the first response should not be automatic model-size escalation. The runtime and benchmark surfaces should instead help answer: - should this gate stay flat or be factorized? - is part of the task actually deterministic? - is the label space unstable? - is the backend family mismatched to the input packet? - should the gate abstain earlier? #### 4A.8 First case-study role Use the memory-strategy decision only as the first end-to-end exercise for the abstract mechanism. The case study should help validate: - mixed symbolic/text/numeric packet design - shape comparison between flat and decomposed gates - backend comparison between classical ML and tiny transformers - packaging of both enum-style and slot-filling outputs It should not be allowed to overfit the library to one memory-management ontology. #### 4A.9 Immediate implementation order - define the AST-first semantic source for RSA contracts - define projection rules from semantic entities into C++ runtime types - extend schema types so both enum and structured-slot gates are first-class - define canonical request/result packet types - define gate-shape metadata and recommendation types - define backend-family contracts, including placeholder RL/bandit support - build the benchmark report format - implement classical ML adapters first - implement transformer adapter second - add packaging and runtime loading - add diagnosis and recommendation APIs on top of benchmark evidence ## Phase 5: Whetstone Integration - Expose WhetstoneDSL gates through the library. - Replace selected MCP bounded decisions with local specialist calls. - Log abstentions and misroutes for retraining and gate redesign. ## Phase 6: Hivemind Integration - Feed gate failure data into Hivemind as entropy evidence. - Use Hivemind to prioritize which high-entropy gates deserve deterministic tooling. - Distinguish "needs a bigger model" from "needs a better tool or schema." ## Exit Criteria For Version 0 - At least one reusable gate schema format is stable. - Both enum-style and bounded structured-slot gates are supported by the contract. - At least three WhetstoneDSL gates are benchmarked across size tiers. - At least one gate has been compared across multiple gate shapes, not just multiple model sizes. - At least one gate has been compared across classical ML and transformer backends under the same contract. - A baseline entropy-to-tier recommendation table exists. - One real application can call the library instead of hard-coding model selection or backend choice.