Files
whetstone_DSL/specialists/HANDOFF-2026-03-28.md

240 lines
10 KiB
Markdown
Raw Normal View History

# Whetstone Specialist Fleet — Handoff 2026-03-28
**Session date:** 2026-03-28
**Status:** Phase 0 complete. Two specialists trained and validated. Clean stop.
---
## What This Directory Is
A fleet of ~800KB transformer specialists that replace bounded-output LLM decisions
in the Whetstone MCP pipeline. Each specialist handles one classification decision,
runs in under a millisecond, and is trained using the WhetstoneAI Fabricate binary
and harness unchanged.
**Origin:** This grew directly from the WhetstoneAI Fabricate RSA experiments
(`/home/bill/Documents/WhetstoneAI_Fabricate/`). The same ~213K-param architecture
that reached 96.3% in-distribution accuracy on domain routing was applied to
Whetstone's bounded taskitem decisions. No new training infrastructure was written.
**Full context:** Read `README.md` and `ROADMAP.md` in this directory before
starting work. They contain the architectural framing, the AST-first design
correction, the 6-tier automatability hierarchy, and the complete next-step plan.
---
## What Was Built This Session
### Two trained specialists
**`verification_type`** — 5-way classifier
- Input: sprint step description text (e.g. `"whetstone_validate_schema MCP tool"`)
- Output: `unit / integration / schema / smoke / docs`
- Training: 978 examples mined from 260 sprint plan files via heuristic labeling
- Eval: 172 held-out examples
- Result: **83.7%** accuracy, grokked at step 1000, stopped at step 3000
- Checkpoint: `/mnt/storage/fabricate_runs/whetstone_verification_type/`
- History: `specialists/runs/verification_type_history.json`
**`confidence_tier`** — 3-way classifier
- Input: structural encoding `"conflicts=N ambiguity=M deps=D prereqs=P queueready=yes/no"`
- Output: `high / medium / low`
- Training: 1020 synthetic examples generated from the formula in `TaskitemConfidenceAmbiguity.h`
- Result: **100%** accuracy, grokked at step 1000, stopped at step 3000
- Checkpoint: `/mnt/storage/fabricate_runs/whetstone_confidence_tier/`
- History: `specialists/runs/confidence_tier_history.json`
### Scripts written
```
specialists/scripts/gen_verification_type_data.py mines sprint plans → TSV
specialists/scripts/gen_confidence_tier_data.py generates synthetic formula data → TSV
specialists/scripts/train_verification_type.sh wraps run_grokking_until.py
specialists/scripts/train_confidence_tier.sh wraps run_grokking_until.py
specialists/scripts/infer.py single/batch/eval-tsv/REPL inference
```
### Inference example
```bash
# From whetstone_DSL root, using fabricate venv:
cd /home/bill/Documents/WhetstoneAI_Fabricate && \
.venv/bin/python3 tools/eval_multiclass.py \
--checkpoint /mnt/storage/fabricate_runs/whetstone_verification_type/checkpoint.bin \
--db /mnt/storage/fabricate_runs/whetstone_verification_type/fabricate.db \
--dataset /home/bill/Documents/CLionProjects/whetstone_DSL/specialists/data/generated/verification_type_eval.tsv \
--labels "unit,integration,schema,smoke,docs"
# → overall: 144/172 = 83.7%
```
---
## Key Technical Facts
**Training data format** (required by Fabricate binary):
```
<integer_label>\t<hops>\t<text>
```
Label is an INTEGER index (0-based) matching the `--labels` flag order.
**String labels in column 1 will crash the binary.** The generators write integers.
**`--labels` flag:** `"unit,integration,schema,smoke,docs"` means label 0=unit,
1=integration, 2=schema, 3=smoke, 4=docs. Order must match the TSV integers.
**Model architecture** (all specialists use this):
`--hidden-dim 128 --layers 1 --heads 1 --seq-len 96 --batch-size 32`
Result: ~213K parameters, ~800KB checkpoint.
**Grokking is normal.** Training logs show "Fresh start" (no prior checkpoint in
output dir), then accuracy jumps at step 1000. This is expected behavior. The
harness stops automatically 3 blocks after grokking.
**Label/checkpoint alignment bug** (from RSA experiments — do not repeat):
If you resume a checkpoint with a changed `--labels` order, the token rows
misalign. Fix with `tools/swap_label_rows.py` in the Fabricate repo.
See `WhetstoneAI_Fabricate/HANDOFF-2026-03-28-RSA.md` for full history.
---
## Important Architectural Decisions Made This Session
### 1. AST-first, not tool-first
Whetstone is an AST-as-source-of-truth system. C++, Python, etc. are projections.
Specialists operate at the AST level — one specialist per decision type covers all
languages. Do NOT train per-language variants of decisions that live in the AST
layer. The projection from AST to language is handled by deterministic grammar rules.
### 2. The 6-tier automatability hierarchy
The current `AutomatabilityAnnotation` in `editor/src/ast/Annotation.h` has 5 tiers.
It is missing the `specialist` tier. The correct hierarchy is:
```
deterministic → formula, zero parameters, always same output
template → deterministic structural composition
specialist → bounded output vocabulary, fuzzy, ~800KB ← THIS PROJECT
slm → open vocabulary generation, ~1-7GB
llm → full generative, ~14B+
human → not yet formalizable
```
Boundary rule: if the output can be expressed as a fixed enum → specialist.
If it requires generating novel tokens → slm or llm.
**`Annotation.h` needs a one-line update** to add `"specialist"` to the strategy
enum before Phase 2.1 training. This is called out in ROADMAP.md.
### 3. Specialist drives template, not replaces it
The specialist and template tiers work as a pipeline, not alternatives:
- Specialist picks which template + fills the bounded slot decisions
- Template deterministically expands those decisions into output
- LLM handles only the residual open-text portions
### 4. The existing AST annotation schema IS the training target vocabulary
`editor/src/ast/Annotation.h` already defines the output schema for most
specialists we need to train. Key fields:
- `AutomatabilityAnnotation.strategy` — the Tacit Codex entropy classifier on nodes
- `IntentAnnotation.category` — validation/transformation/io/coordination/computation/init
- `RiskAnnotation.dependentCount` — downstream caller count (input feature)
- `ComplexityAnnotation.cognitiveComplexity` — 1-10 scale (input feature)
Specialists populate these fields automatically at near-zero cost, replacing
LLM annotation calls.
---
## Exact Next Steps (Phase 1 + Phase 2.1)
### Immediate: Phase 1 pipeline specialists
Four specialists to train for the taskitem pipeline layer. All follow the same
pattern as `verification_type` and `confidence_tier`.
**1.1 — `prereq_op_selector` (4-way)**
Write `scripts/gen_prereq_op_data.py` generating synthetic task descriptions for:
- `standard` (0): [validate-intake, resolve-dependencies] — default path
- `needs_review` (1): [validate-intake, architect-review] — breaking/cross-cutting changes
- `needs_approval` (2): [validate-intake, manual-approval] — experimental/risky tasks
- `full_gates` (3): [validate-intake, architect-review, manual-approval] — both signals
Train: `--labels "standard,needs_review,needs_approval,full_gates"`
**1.2 — `worker_type` (4-way)**
Write `scripts/gen_worker_type_data.py` mining sprint plans with keyword heuristics:
- `implementer` (0): implement/add/build/write/create
- `reviewer` (1): review/audit/verify/check/validate
- `architect` (2): design/architect/plan/structure/refactor
- `qa` (3): test/coverage/regression/integration summary
Train: `--labels "implementer,reviewer,architect,qa"`
**1.3 — `gap_class` (5-way)**
Write `scripts/gen_gap_class_data.py` with synthetic taskitem issue descriptions.
Classes: under_constrained / environment_gap / planning_gap / cross_project_gap / generator_capability_gap
**1.4 — `escalation_predictor` (binary)**
Write `scripts/gen_escalation_data.py` from formula: `escalate = confidence < 60 OR ambiguity > 0`.
Expected 100% accuracy like `confidence_tier`.
Train: `--labels "proceed,escalate"`
### After Phase 1: Phase 2.1 — `automatability_strategy` (6-way)
**First:** Add `"specialist"` to strategy enum in `editor/src/ast/Annotation.h`.
Write `scripts/gen_automatability_data.py` using the 93 existing `editor/src/*.h`
headers as labeled examples. Each header encodes the ground-truth tier:
- Pure formula classes → `deterministic`
- Schema-to-X converters → `template`
- Classifiers/scorers with bounded output → `specialist`
- Text generators with constrained vocabulary → `slm`
- Open-ended generators → `llm`
Train: `--labels "deterministic,template,specialist,slm,llm,human"`
This is the most architecturally significant specialist — once trained, every node
in the AST can be tagged with its automatability tier without an LLM call.
---
## What Is NOT Done / Known Gaps
- `infer.py` is a standalone inference tool but not yet wired into any MCP tool
- No dispatch server yet (`scripts/dispatch.py` — Phase 4.1 in roadmap)
- `Annotation.h` still has the old 5-tier enum (needs `"specialist"` added)
- The `verification_type` specialist has 16% error rate concentrated at the
schema/unit/integration boundary — the heuristic labeling is imperfect there.
Acceptable for now; fix by reviewing ~50 misclassified examples if accuracy matters.
- JSONL pipeline run records (18,519) are too homogeneous for text-based specialists.
Sprint plans are the right source. Do not try to use JSONL for text classifiers.
---
## Connection to WhetstoneAI Fabricate
The Fabricate binary and harness are consumed as-is. Do not modify them.
```
Binary: /home/bill/Documents/WhetstoneAI_Fabricate/build/whetstone_fabricate
Harness: /home/bill/Documents/WhetstoneAI_Fabricate/run_grokking_until.py
Eval: /home/bill/Documents/WhetstoneAI_Fabricate/tools/eval_multiclass.py
Infer: /home/bill/Documents/WhetstoneAI_Fabricate/tools/eval_with_vocab.py
```
The Fabricate venv at `/home/bill/Documents/WhetstoneAI_Fabricate/.venv` has all
required packages. Use `.venv/bin/python3` for all training and eval calls.
---
## Connection to Tacit Codex
The specialist fleet is the concrete implementation of the Tacit Codex program:
each trained specialist is a named, deployed piece of collapsed tacit knowledge.
The `AutomatabilityAnnotation.strategy` field, once populated by the Phase 2.1
specialist, directly encodes the entropy tier of every AST node — which is the
per-node Tacit Codex entropy ledger entry.
The same pattern (hard constraints + fuzzy contextual decisions) also applies to
WHIMP (interaction design) with the 722-metric DuckDB as the input feature vocabulary.
That connection is documented in the session conversation but not yet in any file.