Files
whetstone_DSL/specialists/HANDOFF-2026-03-31.md
Bill 3fc9deac5b specialists: switch training to PyTorch, add capacity sweep infrastructure
- train_specialist_pt.py: pure PyTorch trainer replacing Fabricate binary.
  No SQLite/WAL — saves checkpoint.pt + history.json only. Fabricate was
  causing D-state IO blocking on the HDD due to continuous WAL writes.
- eval_specialist_pt.py: PyTorch eval with confusion matrix, threshold
  analysis, guardrail assessment, and deployment verdict.
- capacity_sweep.py: updated to call PyTorch scripts; TIERS now include
  heads field (4/6/8) since PyTorch supports multi-head unlike Fabricate.
- prereq_op_binary_split.py: evaluate_combined rewritten in PyTorch.
- Performance fix in train_specialist_pt.py: load training data onto GPU
  once and sample via torch.randint instead of DataLoader (eliminates
  Python/CPU overhead on tiny datasets). Fix applied, not yet timed.

Next session: time the fix, then run sweep_all_gates.sh --pilot-only.
See HANDOFF-2026-03-31.md for full context.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-31 22:55:51 -06:00

4.4 KiB
Raw Permalink Blame History

Specialist Fleet Handoff — 2026-03-31

Where We Are

Switched training infrastructure from Fabricate (C++ binary + SQLite) to pure PyTorch. Fabricate was causing uninterruptible D-state IO blocking because /mnt/storage is a spinning HDD and Fabricate writes continuous SQLite WAL data (~100MB/min) even with morphisms disabled.

What Was Built This Session

New Files

  • specialists/scripts/train_specialist_pt.py — PyTorch training script (replaces run_grokking_until.py)
  • specialists/eval/eval_specialist_pt.py — PyTorch eval script (replaces eval_specialist.py)

Updated Files

  • specialists/scripts/capacity_sweep.py — now calls train_specialist_pt.py + eval_specialist_pt.py
  • specialists/scripts/prereq_op_binary_split.py — same, evaluate_combined rewritten in PyTorch
  • Both scripts now use checkpoint.pt instead of checkpoint.bin + fabricate.db
  • TIERS updated: heads field added (4/6/8 for small+/medium/large) — PyTorch supports multi-head unlike Fabricate

PyTorch training script details

  • No SQLite, no WAL. Saves checkpoint.pt + history.json only.
  • History JSON format compatible with capacity_sweep.py: {"history": [{"step": N, "overall": float, "loss": float}]}
  • Same CLI interface as run_grokking_until.py for drop-in use by sweep scripts
  • Grokking detection + early stop implemented (same logic: loss < threshold OR acc jump ≥ threshold)
  • Checkpoint includes: model weights, optimizer state, vocab, config, labels, step count

Performance Issue — PARTIALLY FIXED

Root cause: DataLoader restarts ~833×/1000 steps when dataset has 300 examples + batch_size=256. Each restart has Python overhead → 38s per 1000-step block (expected ~4s).

Fix applied (untimed): Load all training data onto GPU once as a tensor, sample via torch.randint each step instead of DataLoader. This eliminates all Python/CPU overhead per step.

all_x = torch.stack([row[0] for row in train_ds]).to(device)
all_y = torch.tensor([row[1] for row in train_ds], ..., device=device)
# per step:
idx = torch.randint(0, n_train, (batch_size,), device=device)
x, y = all_x[idx], all_y[idx]

This fix is in the code but was NOT timed before session ended.

First Task Next Session

  1. Time the fixed training script:
cd /home/bill/Documents/CLionProjects/whetstone_DSL
time /home/bill/Documents/WhetstoneAI_Fabricate/.venv/bin/python3 \
  specialists/scripts/train_specialist_pt.py \
  --dataset specialists/data/generated/worker_type_train.tsv \
  --eval-dataset specialists/data/generated/worker_type_eval.tsv \
  --labels "implementer,reviewer,architect,qa" \
  --out-dir /tmp/pt_timing_test \
  --hidden-dim 256 --layers 2 --heads 4 \
  --lr 0.001 --weight-decay 0.01 \
  --max-steps 3000 --batch-size 256 \
  --stop-after-grokking-blocks 99 \
  --reset

Expected: ~4-8 seconds per 1000-step block (was 38s before fix), so ~12-25s for 3000 steps. If correct, pilot sweep (all 4 gates × 3 tiers) should complete in ~5-10 minutes.

  1. If timing is good, run:
bash specialists/scripts/sweep_all_gates.sh --pilot-only
  1. Analyze pilot results in specialists/eval/results/summary.json, decide which gates need full training vs already good enough.

Current Baseline Accuracies (small Fabricate models, unchanged)

  • worker_type: 88.5% (implementer residual-category issue, 64% implementer acc)
  • verification_type: 84.9% (schema→integration confusion, capacity limited)
  • prereq_op: 69.4% (standard residual-category, over-fires approval gate)
  • automatability: 69.6% (deterministic 0% — pure residual-category collapse)

Target: 98-99% for all gates. Larger PyTorch models (no multi-head constraint) should help.

Key Architecture Notes

  • Fabricate constraint: gpu_transformer engine only supports heads=1
  • PyTorch: any heads where hidden_dim % heads == 0 — we now use 4/6/8
  • Run dirs are still on /mnt/storage/fabricate_runs/ — this is fine, .pt files are tiny (few MB) and written infrequently (once per 1000 steps), not continuously
  • The old whetstone_prereq_op_large run dir on the HDD may still have a D-state Fabricate process associated. Check pgrep fabricate at session start; kill if present.

Venv

All scripts use: /home/bill/Documents/WhetstoneAI_Fabricate/.venv/bin/python3 PyTorch 2.10+cu128 is installed there. CUDA is available (RTX 3060).