specialists: switch training to PyTorch, add capacity sweep infrastructure

- train_specialist_pt.py: pure PyTorch trainer replacing Fabricate binary.
  No SQLite/WAL — saves checkpoint.pt + history.json only. Fabricate was
  causing D-state IO blocking on the HDD due to continuous WAL writes.
- eval_specialist_pt.py: PyTorch eval with confusion matrix, threshold
  analysis, guardrail assessment, and deployment verdict.
- capacity_sweep.py: updated to call PyTorch scripts; TIERS now include
  heads field (4/6/8) since PyTorch supports multi-head unlike Fabricate.
- prereq_op_binary_split.py: evaluate_combined rewritten in PyTorch.
- Performance fix in train_specialist_pt.py: load training data onto GPU
  once and sample via torch.randint instead of DataLoader (eliminates
  Python/CPU overhead on tiny datasets). Fix applied, not yet timed.

Next session: time the fix, then run sweep_all_gates.sh --pilot-only.
See HANDOFF-2026-03-31.md for full context.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Bill
2026-03-31 22:55:51 -06:00
parent 410bb2d78c
commit 3fc9deac5b
61 changed files with 9168 additions and 0 deletions

View File

@@ -0,0 +1,197 @@
# Whetstone Specialist Fleet — Architecture
**Date:** 2026-03-28
**Status:** Phase 1 — first specialist training
---
## The Core Idea
The current Whetstone pipeline makes every bounded decision with one large LLM call.
Each call is ~seconds, sequential, and expensive. Many of these decisions are
bounded-output classifications — the exact problem that WhetstoneAI Fabricate's
~213K-parameter (~800KB) specialists solve in microseconds.
**Hypothesis:** Replace bounded-output LLM decisions in the Whetstone taskitem
pipeline with a fleet of micro-specialists, each ~800KB, running in parallel.
### Why It Works
The WhetstoneAI RSA experiment demonstrated:
- 96.3% in-distribution accuracy on domain classification
- 76.7% OOD accuracy after replay hardening (up from 44.8%)
- Inference time: microseconds per query on GPU, 15ms on CPU (numpy)
- Model size: ~213K params, ~800KB per checkpoint
- RTX 3060 (12GB VRAM): can hold ~14,000 specialists resident simultaneously
The Whetstone taskitem pipeline already has:
- Bounded enum decision points (verification type, prerequisite ops, worker type...)
- Recorded task completions (18,519 pipeline runs)
- Concurrency/dependency architecture (lock-based pool with parallel workers)
This means the specialists can map 1:1 onto the pipeline's decision points and
run in parallel with the existing task dependency graph.
---
## Decision Points Amenable to Specialists
### Tier 1: Text-Input Classifiers (primary target)
These decisions receive a task description and output a bounded choice.
They currently run inside the LLM and are the highest-value targets.
| Specialist | Input | Output | Classes |
|-----------|-------|--------|---------|
| `verification_type` | task step description | verification type | unit / integration / schema / smoke / docs |
| `prereq_op_selector` | task description | prerequisite op set | validate-intake / resolve-dependencies / architect-review / manual-approval |
| `worker_type` | task description | worker role | implementer / reviewer / architect / qa |
| `gap_class` | taskitem issues list | gap classification | under_constrained / env_gap / planning_gap / cross_project / generator_cap |
### Tier 2: Structural Specialists (formula approximators)
These learn the deterministic scoring formulas — useful for ultra-fast pre-screening
without spinning up the full validator.
| Specialist | Input | Output | Notes |
|-----------|-------|--------|-------|
| `confidence_tier` | (conflict_count, ambiguity_count, deps, prereqs) | high / medium / low | Learns the deduction formula |
| `escalation_predictor` | (confidence, ambiguity_count) | escalate: bool | Binary: confidence < 60 OR ambiguity > 0 |
| `quality_score_tier` | task features | pass / warn / fail | Learns SelfContainmentScorer cutoffs |
### Tier 3: Future Specialists (need data collection)
| Specialist | Input | Output | Blocker |
|-----------|-------|--------|---------|
| `library_selector` | task intent + available libs | selected library | No labeled data yet |
| `milestone_count` | requirement count + conflicts | 1 / 2 / 3 / 4+ milestones | Need more diverse pipeline runs |
| `task_position_to_dependency` | task index + total count | has_dep: bool | Imbalanced data (97% no-dep) |
---
## Architecture: Fleet Model
```
┌─────────────────────────────────────────────────────────────────┐
│ Whetstone MCP Request │
│ (generate_taskitems, validate_taskitem, ...) │
└────────────────────────┬────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────┐
│ Specialist Dispatcher (Python: infer.py) │
│ │
│ Task text ──► verification_type_specialist ──► unit │
│ Task text ──► prereq_op_specialist ──► [validate, ...] │
│ Task text ──► worker_type_specialist ──► implementer │
│ Features ──► confidence_tier_specialist ──► high │
│ │
│ All specialists run in parallel via thread pool │
└──────────────┬──────────────────────────────────────────────────┘
│ aggregated results
┌─────────────────────────────────────────────────────────────────┐
│ Whetstone C++ Runtime │
│ (enriches taskitem fields with specialist predictions) │
└─────────────────────────────────────────────────────────────────┘
```
### Specialist Model Spec (all identical architecture)
```
Architecture: Transformer (D=128, L=1, heads=1, seq_len=96)
Parameters: ~213K
Size: ~832KB per checkpoint
Inference: <1ms CPU (numpy), ~50µs GPU
Training: 1K5K steps, grokking detection, weight decay
```
Reuses the WhetstoneAI Fabricate binary (`whetstone_fabricate`) and training harness
(`run_grokking_until.py`) without modification.
---
## Training Pipeline
```
sprint_plans/*.md
scripts/gen_verification_type_data.py ← TSV extractor (one per specialist)
data/generated/verification_type_train.tsv (label<TAB>0<TAB>text)
data/generated/verification_type_eval.tsv
scripts/train_verification_type.sh ← wraps run_grokking_until.py
/mnt/storage/fabricate_runs/whetstone_verification_type/
checkpoint.bin (weights + optional Adam state)
fabricate.db (vocabulary)
specialists/scripts/infer.py ← Python inference wrapper
Whetstone MCP tool integration (future)
```
---
## Data Sources
### Current (18,519 pipeline run records)
`training_data/lora/taskitem_pipeline_runs.jsonl`
**Limitation:** Almost entirely homogeneous (2 task types repeated, no text diversity).
Useful for structural/formula specialists (Tier 2). NOT useful for text-based classifiers.
### Primary for Text Specialists
`sprint*_plan.md` files (260 files, ~2,600 labeled step descriptions)
Each step description has an implicit verification type based on its content:
- `### Step NNN: whetstone_foo MCP tool` → integration
- `### Step NNN: Schema for X` → schema
- `### Step NNN: Sprint N integration summary + regression` → smoke
- `### Step NNN: X documentation packet` → docs
- Everything else → unit
### Future Data Collection
Add richer recording to `run_sprint_taskitem_pipeline.sh`:
- Capture per-step verification type decisions
- Record library selection choices
- Record milestone count decisions
---
## Current Status
| Specialist | Data Source | Training Status | Accuracy |
|-----------|------------|-----------------|---------|
| `verification_type` | sprint plans (978 train, 172 eval) | **TRAINING** | TBD |
| All others | — | planned | — |
---
## Integration Plan (Future)
1. **Phase 1 (now):** Train verification_type specialist, validate accuracy
2. **Phase 2:** Wire into `whetstone_validate_taskitem` as a pre-prediction step
3. **Phase 3:** Train prereq_op and confidence_tier specialists
4. **Phase 4:** Build Python dispatcher that runs all specialists in parallel
5. **Phase 5:** Expose as new MCP tool `whetstone_specialist_predict`
6. **Phase 6:** Replace LLM calls for bounded decisions with specialist fleet
---
## Connection to WhetstoneAI Fabricate
This project reuses the Fabricate infrastructure directly:
- Binary: `/home/bill/Documents/WhetstoneAI_Fabricate/build/whetstone_fabricate`
- Harness: `/home/bill/Documents/WhetstoneAI_Fabricate/run_grokking_until.py`
- Eval tools: `/home/bill/Documents/WhetstoneAI_Fabricate/tools/eval_multiclass.py`
- Inference: `/home/bill/Documents/WhetstoneAI_Fabricate/tools/eval_with_vocab.py`
No code changes to Fabricate are needed. It is consumed as a training and inference substrate.