- train_specialist_pt.py: pure PyTorch trainer replacing Fabricate binary. No SQLite/WAL — saves checkpoint.pt + history.json only. Fabricate was causing D-state IO blocking on the HDD due to continuous WAL writes. - eval_specialist_pt.py: PyTorch eval with confusion matrix, threshold analysis, guardrail assessment, and deployment verdict. - capacity_sweep.py: updated to call PyTorch scripts; TIERS now include heads field (4/6/8) since PyTorch supports multi-head unlike Fabricate. - prereq_op_binary_split.py: evaluate_combined rewritten in PyTorch. - Performance fix in train_specialist_pt.py: load training data onto GPU once and sample via torch.randint instead of DataLoader (eliminates Python/CPU overhead on tiny datasets). Fix applied, not yet timed. Next session: time the fix, then run sweep_all_gates.sh --pilot-only. See HANDOFF-2026-03-31.md for full context. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
8.7 KiB
Whetstone Specialist Fleet — Architecture
Date: 2026-03-28 Status: Phase 1 — first specialist training
The Core Idea
The current Whetstone pipeline makes every bounded decision with one large LLM call. Each call is ~seconds, sequential, and expensive. Many of these decisions are bounded-output classifications — the exact problem that WhetstoneAI Fabricate's ~213K-parameter (~800KB) specialists solve in microseconds.
Hypothesis: Replace bounded-output LLM decisions in the Whetstone taskitem pipeline with a fleet of micro-specialists, each ~800KB, running in parallel.
Why It Works
The WhetstoneAI RSA experiment demonstrated:
- 96.3% in-distribution accuracy on domain classification
- 76.7% OOD accuracy after replay hardening (up from 44.8%)
- Inference time: microseconds per query on GPU, 1–5ms on CPU (numpy)
- Model size: ~213K params, ~800KB per checkpoint
- RTX 3060 (12GB VRAM): can hold ~14,000 specialists resident simultaneously
The Whetstone taskitem pipeline already has:
- Bounded enum decision points (verification type, prerequisite ops, worker type...)
- Recorded task completions (18,519 pipeline runs)
- Concurrency/dependency architecture (lock-based pool with parallel workers)
This means the specialists can map 1:1 onto the pipeline's decision points and run in parallel with the existing task dependency graph.
Decision Points Amenable to Specialists
Tier 1: Text-Input Classifiers (primary target)
These decisions receive a task description and output a bounded choice. They currently run inside the LLM and are the highest-value targets.
| Specialist | Input | Output | Classes |
|---|---|---|---|
verification_type |
task step description | verification type | unit / integration / schema / smoke / docs |
prereq_op_selector |
task description | prerequisite op set | validate-intake / resolve-dependencies / architect-review / manual-approval |
worker_type |
task description | worker role | implementer / reviewer / architect / qa |
gap_class |
taskitem issues list | gap classification | under_constrained / env_gap / planning_gap / cross_project / generator_cap |
Tier 2: Structural Specialists (formula approximators)
These learn the deterministic scoring formulas — useful for ultra-fast pre-screening without spinning up the full validator.
| Specialist | Input | Output | Notes |
|---|---|---|---|
confidence_tier |
(conflict_count, ambiguity_count, deps, prereqs) | high / medium / low | Learns the deduction formula |
escalation_predictor |
(confidence, ambiguity_count) | escalate: bool | Binary: confidence < 60 OR ambiguity > 0 |
quality_score_tier |
task features | pass / warn / fail | Learns SelfContainmentScorer cutoffs |
Tier 3: Future Specialists (need data collection)
| Specialist | Input | Output | Blocker |
|---|---|---|---|
library_selector |
task intent + available libs | selected library | No labeled data yet |
milestone_count |
requirement count + conflicts | 1 / 2 / 3 / 4+ milestones | Need more diverse pipeline runs |
task_position_to_dependency |
task index + total count | has_dep: bool | Imbalanced data (97% no-dep) |
Architecture: Fleet Model
┌─────────────────────────────────────────────────────────────────┐
│ Whetstone MCP Request │
│ (generate_taskitems, validate_taskitem, ...) │
└────────────────────────┬────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ Specialist Dispatcher (Python: infer.py) │
│ │
│ Task text ──► verification_type_specialist ──► unit │
│ Task text ──► prereq_op_specialist ──► [validate, ...] │
│ Task text ──► worker_type_specialist ──► implementer │
│ Features ──► confidence_tier_specialist ──► high │
│ │
│ All specialists run in parallel via thread pool │
└──────────────┬──────────────────────────────────────────────────┘
│ aggregated results
▼
┌─────────────────────────────────────────────────────────────────┐
│ Whetstone C++ Runtime │
│ (enriches taskitem fields with specialist predictions) │
└─────────────────────────────────────────────────────────────────┘
Specialist Model Spec (all identical architecture)
Architecture: Transformer (D=128, L=1, heads=1, seq_len=96)
Parameters: ~213K
Size: ~832KB per checkpoint
Inference: <1ms CPU (numpy), ~50µs GPU
Training: 1K–5K steps, grokking detection, weight decay
Reuses the WhetstoneAI Fabricate binary (whetstone_fabricate) and training harness
(run_grokking_until.py) without modification.
Training Pipeline
sprint_plans/*.md
│
▼
scripts/gen_verification_type_data.py ← TSV extractor (one per specialist)
│
▼
data/generated/verification_type_train.tsv (label<TAB>0<TAB>text)
data/generated/verification_type_eval.tsv
│
▼
scripts/train_verification_type.sh ← wraps run_grokking_until.py
│
▼
/mnt/storage/fabricate_runs/whetstone_verification_type/
checkpoint.bin (weights + optional Adam state)
fabricate.db (vocabulary)
│
▼
specialists/scripts/infer.py ← Python inference wrapper
│
▼
Whetstone MCP tool integration (future)
Data Sources
Current (18,519 pipeline run records)
training_data/lora/taskitem_pipeline_runs.jsonl
Limitation: Almost entirely homogeneous (2 task types repeated, no text diversity). Useful for structural/formula specialists (Tier 2). NOT useful for text-based classifiers.
Primary for Text Specialists
sprint*_plan.md files (260 files, ~2,600 labeled step descriptions)
Each step description has an implicit verification type based on its content:
### Step NNN: whetstone_foo MCP tool→ integration### Step NNN: Schema for X→ schema### Step NNN: Sprint N integration summary + regression→ smoke### Step NNN: X documentation packet→ docs- Everything else → unit
Future Data Collection
Add richer recording to run_sprint_taskitem_pipeline.sh:
- Capture per-step verification type decisions
- Record library selection choices
- Record milestone count decisions
Current Status
| Specialist | Data Source | Training Status | Accuracy |
|---|---|---|---|
verification_type |
sprint plans (978 train, 172 eval) | TRAINING | TBD |
| All others | — | planned | — |
Integration Plan (Future)
- Phase 1 (now): Train verification_type specialist, validate accuracy
- Phase 2: Wire into
whetstone_validate_taskitemas a pre-prediction step - Phase 3: Train prereq_op and confidence_tier specialists
- Phase 4: Build Python dispatcher that runs all specialists in parallel
- Phase 5: Expose as new MCP tool
whetstone_specialist_predict - Phase 6: Replace LLM calls for bounded decisions with specialist fleet
Connection to WhetstoneAI Fabricate
This project reuses the Fabricate infrastructure directly:
- Binary:
/home/bill/Documents/WhetstoneAI_Fabricate/build/whetstone_fabricate - Harness:
/home/bill/Documents/WhetstoneAI_Fabricate/run_grokking_until.py - Eval tools:
/home/bill/Documents/WhetstoneAI_Fabricate/tools/eval_multiclass.py - Inference:
/home/bill/Documents/WhetstoneAI_Fabricate/tools/eval_with_vocab.py
No code changes to Fabricate are needed. It is consumed as a training and inference substrate.