- train_specialist_pt.py: pure PyTorch trainer replacing Fabricate binary. No SQLite/WAL — saves checkpoint.pt + history.json only. Fabricate was causing D-state IO blocking on the HDD due to continuous WAL writes. - eval_specialist_pt.py: PyTorch eval with confusion matrix, threshold analysis, guardrail assessment, and deployment verdict. - capacity_sweep.py: updated to call PyTorch scripts; TIERS now include heads field (4/6/8) since PyTorch supports multi-head unlike Fabricate. - prereq_op_binary_split.py: evaluate_combined rewritten in PyTorch. - Performance fix in train_specialist_pt.py: load training data onto GPU once and sample via torch.randint instead of DataLoader (eliminates Python/CPU overhead on tiny datasets). Fix applied, not yet timed. Next session: time the fix, then run sweep_all_gates.sh --pilot-only. See HANDOFF-2026-03-31.md for full context. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
198 lines
8.7 KiB
Markdown
198 lines
8.7 KiB
Markdown
# Whetstone Specialist Fleet — Architecture
|
||
|
||
**Date:** 2026-03-28
|
||
**Status:** Phase 1 — first specialist training
|
||
|
||
---
|
||
|
||
## The Core Idea
|
||
|
||
The current Whetstone pipeline makes every bounded decision with one large LLM call.
|
||
Each call is ~seconds, sequential, and expensive. Many of these decisions are
|
||
bounded-output classifications — the exact problem that WhetstoneAI Fabricate's
|
||
~213K-parameter (~800KB) specialists solve in microseconds.
|
||
|
||
**Hypothesis:** Replace bounded-output LLM decisions in the Whetstone taskitem
|
||
pipeline with a fleet of micro-specialists, each ~800KB, running in parallel.
|
||
|
||
### Why It Works
|
||
|
||
The WhetstoneAI RSA experiment demonstrated:
|
||
- 96.3% in-distribution accuracy on domain classification
|
||
- 76.7% OOD accuracy after replay hardening (up from 44.8%)
|
||
- Inference time: microseconds per query on GPU, 1–5ms on CPU (numpy)
|
||
- Model size: ~213K params, ~800KB per checkpoint
|
||
- RTX 3060 (12GB VRAM): can hold ~14,000 specialists resident simultaneously
|
||
|
||
The Whetstone taskitem pipeline already has:
|
||
- Bounded enum decision points (verification type, prerequisite ops, worker type...)
|
||
- Recorded task completions (18,519 pipeline runs)
|
||
- Concurrency/dependency architecture (lock-based pool with parallel workers)
|
||
|
||
This means the specialists can map 1:1 onto the pipeline's decision points and
|
||
run in parallel with the existing task dependency graph.
|
||
|
||
---
|
||
|
||
## Decision Points Amenable to Specialists
|
||
|
||
### Tier 1: Text-Input Classifiers (primary target)
|
||
|
||
These decisions receive a task description and output a bounded choice.
|
||
They currently run inside the LLM and are the highest-value targets.
|
||
|
||
| Specialist | Input | Output | Classes |
|
||
|-----------|-------|--------|---------|
|
||
| `verification_type` | task step description | verification type | unit / integration / schema / smoke / docs |
|
||
| `prereq_op_selector` | task description | prerequisite op set | validate-intake / resolve-dependencies / architect-review / manual-approval |
|
||
| `worker_type` | task description | worker role | implementer / reviewer / architect / qa |
|
||
| `gap_class` | taskitem issues list | gap classification | under_constrained / env_gap / planning_gap / cross_project / generator_cap |
|
||
|
||
### Tier 2: Structural Specialists (formula approximators)
|
||
|
||
These learn the deterministic scoring formulas — useful for ultra-fast pre-screening
|
||
without spinning up the full validator.
|
||
|
||
| Specialist | Input | Output | Notes |
|
||
|-----------|-------|--------|-------|
|
||
| `confidence_tier` | (conflict_count, ambiguity_count, deps, prereqs) | high / medium / low | Learns the deduction formula |
|
||
| `escalation_predictor` | (confidence, ambiguity_count) | escalate: bool | Binary: confidence < 60 OR ambiguity > 0 |
|
||
| `quality_score_tier` | task features | pass / warn / fail | Learns SelfContainmentScorer cutoffs |
|
||
|
||
### Tier 3: Future Specialists (need data collection)
|
||
|
||
| Specialist | Input | Output | Blocker |
|
||
|-----------|-------|--------|---------|
|
||
| `library_selector` | task intent + available libs | selected library | No labeled data yet |
|
||
| `milestone_count` | requirement count + conflicts | 1 / 2 / 3 / 4+ milestones | Need more diverse pipeline runs |
|
||
| `task_position_to_dependency` | task index + total count | has_dep: bool | Imbalanced data (97% no-dep) |
|
||
|
||
---
|
||
|
||
## Architecture: Fleet Model
|
||
|
||
```
|
||
┌─────────────────────────────────────────────────────────────────┐
|
||
│ Whetstone MCP Request │
|
||
│ (generate_taskitems, validate_taskitem, ...) │
|
||
└────────────────────────┬────────────────────────────────────────┘
|
||
│
|
||
▼
|
||
┌─────────────────────────────────────────────────────────────────┐
|
||
│ Specialist Dispatcher (Python: infer.py) │
|
||
│ │
|
||
│ Task text ──► verification_type_specialist ──► unit │
|
||
│ Task text ──► prereq_op_specialist ──► [validate, ...] │
|
||
│ Task text ──► worker_type_specialist ──► implementer │
|
||
│ Features ──► confidence_tier_specialist ──► high │
|
||
│ │
|
||
│ All specialists run in parallel via thread pool │
|
||
└──────────────┬──────────────────────────────────────────────────┘
|
||
│ aggregated results
|
||
▼
|
||
┌─────────────────────────────────────────────────────────────────┐
|
||
│ Whetstone C++ Runtime │
|
||
│ (enriches taskitem fields with specialist predictions) │
|
||
└─────────────────────────────────────────────────────────────────┘
|
||
```
|
||
|
||
### Specialist Model Spec (all identical architecture)
|
||
|
||
```
|
||
Architecture: Transformer (D=128, L=1, heads=1, seq_len=96)
|
||
Parameters: ~213K
|
||
Size: ~832KB per checkpoint
|
||
Inference: <1ms CPU (numpy), ~50µs GPU
|
||
Training: 1K–5K steps, grokking detection, weight decay
|
||
```
|
||
|
||
Reuses the WhetstoneAI Fabricate binary (`whetstone_fabricate`) and training harness
|
||
(`run_grokking_until.py`) without modification.
|
||
|
||
---
|
||
|
||
## Training Pipeline
|
||
|
||
```
|
||
sprint_plans/*.md
|
||
│
|
||
▼
|
||
scripts/gen_verification_type_data.py ← TSV extractor (one per specialist)
|
||
│
|
||
▼
|
||
data/generated/verification_type_train.tsv (label<TAB>0<TAB>text)
|
||
data/generated/verification_type_eval.tsv
|
||
│
|
||
▼
|
||
scripts/train_verification_type.sh ← wraps run_grokking_until.py
|
||
│
|
||
▼
|
||
/mnt/storage/fabricate_runs/whetstone_verification_type/
|
||
checkpoint.bin (weights + optional Adam state)
|
||
fabricate.db (vocabulary)
|
||
│
|
||
▼
|
||
specialists/scripts/infer.py ← Python inference wrapper
|
||
│
|
||
▼
|
||
Whetstone MCP tool integration (future)
|
||
```
|
||
|
||
---
|
||
|
||
## Data Sources
|
||
|
||
### Current (18,519 pipeline run records)
|
||
`training_data/lora/taskitem_pipeline_runs.jsonl`
|
||
|
||
**Limitation:** Almost entirely homogeneous (2 task types repeated, no text diversity).
|
||
Useful for structural/formula specialists (Tier 2). NOT useful for text-based classifiers.
|
||
|
||
### Primary for Text Specialists
|
||
`sprint*_plan.md` files (260 files, ~2,600 labeled step descriptions)
|
||
|
||
Each step description has an implicit verification type based on its content:
|
||
- `### Step NNN: whetstone_foo MCP tool` → integration
|
||
- `### Step NNN: Schema for X` → schema
|
||
- `### Step NNN: Sprint N integration summary + regression` → smoke
|
||
- `### Step NNN: X documentation packet` → docs
|
||
- Everything else → unit
|
||
|
||
### Future Data Collection
|
||
Add richer recording to `run_sprint_taskitem_pipeline.sh`:
|
||
- Capture per-step verification type decisions
|
||
- Record library selection choices
|
||
- Record milestone count decisions
|
||
|
||
---
|
||
|
||
## Current Status
|
||
|
||
| Specialist | Data Source | Training Status | Accuracy |
|
||
|-----------|------------|-----------------|---------|
|
||
| `verification_type` | sprint plans (978 train, 172 eval) | **TRAINING** | TBD |
|
||
| All others | — | planned | — |
|
||
|
||
---
|
||
|
||
## Integration Plan (Future)
|
||
|
||
1. **Phase 1 (now):** Train verification_type specialist, validate accuracy
|
||
2. **Phase 2:** Wire into `whetstone_validate_taskitem` as a pre-prediction step
|
||
3. **Phase 3:** Train prereq_op and confidence_tier specialists
|
||
4. **Phase 4:** Build Python dispatcher that runs all specialists in parallel
|
||
5. **Phase 5:** Expose as new MCP tool `whetstone_specialist_predict`
|
||
6. **Phase 6:** Replace LLM calls for bounded decisions with specialist fleet
|
||
|
||
---
|
||
|
||
## Connection to WhetstoneAI Fabricate
|
||
|
||
This project reuses the Fabricate infrastructure directly:
|
||
- Binary: `/home/bill/Documents/WhetstoneAI_Fabricate/build/whetstone_fabricate`
|
||
- Harness: `/home/bill/Documents/WhetstoneAI_Fabricate/run_grokking_until.py`
|
||
- Eval tools: `/home/bill/Documents/WhetstoneAI_Fabricate/tools/eval_multiclass.py`
|
||
- Inference: `/home/bill/Documents/WhetstoneAI_Fabricate/tools/eval_with_vocab.py`
|
||
|
||
No code changes to Fabricate are needed. It is consumed as a training and inference substrate.
|