# Handoff: RSA Session 2 — 2026-04-12 ## What Was Done ### Extraction and dataset split - Built `tools/split_prereq_op_dataset.py` - splits extracted gate rows into accepted / rejected / schema-drift buckets - for `prereq_op_selector` only - Ran `extract_gate_rows.py` over 1000 runs (from 50-run slice) - 1297 accepted rows, 24 rejected (all schema-drift) - outputs at `semantic/prereq_op_dataset_1000/` - Built `tools/gen_prereq_op_rsa_data.py` - converts accepted rows to Fabricate TSV format - outputs at `semantic/prereq_op_tsv/` ### Key findings from dataset analysis **`needs_validate_intake` is always True** (1297/1297 accepted rows). It is a constant, not a decision surface. Drop it from the label space. **Effective gate is two binary classifiers:** - `needs_resolve_dependencies`: 997/1297 = 76.9% positive - `needs_architect_review`: 257/1297 = 19.8% positive **Text features are useless from run corpus.** Only 17 unique texts across 1297 rows — the corpus uses ~18 canonical intake spec templates. Every task generated from the same intake shares the same context text. The discriminative signal (step description) is not stored in run artifacts. **Schema-drift variants found:** - tool-ID drift: `whetstone_generate_taskitems` etc. in `prerequisiteOps` (12 rows) - `manual-approval` drift: undocumented op name in `prerequisiteOps` (12 rows) Both are salvageable (valid bits are correct), recommended to hard-filter for now. ### Training baseline Sprint plan data (204 examples, 36 eval) split into two binary TSVs and trained: | Specialist | best_acc | steps | |------------|----------|-------| | `prereq_review` (needs architect_review) | 75.0% | 4000 | | `prereq_approval` (needs manual_approval) | 69.4% | 4000 | These are weak baselines — small dataset, heuristic labels, wrong gate framing (approval ≠ resolve_dependencies). Not production-ready. ## Architectural Insights Captured (Most Important Part of This Session) ### 1. Intake contract The original design intention is that external tools (HiveMind, gleaner, etc.) submit project descriptions in a **structured contract format** that explicitly carries decision-relevant fields. The current corpus predates this contract. With the contract, every project submission is a varied, labeled training example with known discriminative fields. Without it, the intake produces ~18 template outputs regardless of project. Reference: `project.json` manifests are the closest existing contract artifact. The `llm_touchpoints` array with `determinism_score` is the RSA entropy signal. ### 2. Taskitem as self-contained work unit A taskitem is a concurrent work unit. By design it must include all information the specialist needs. The specialist does not reach back to the original spec. **This means the taskitem schema must carry discriminative RSA fields.** `generate_taskitems` populates them; the RSA gate reads them from the taskitem. The current schema lacks these fields — that is why the run corpus has weak text. Adding RSA-relevant fields to the taskitem schema is the **prerequisite** for collecting real training data at scale. ### 3. Contract can collapse RSA gates to deterministic If the contract carries `architectural_surface`, `cross_component_deps`, `security_sensitive` etc. explicitly, some gates that currently require ML become deterministic rules. The RSA entropy threshold shifts as the contract gets richer. ## What Is Recorded Where | Finding | Location | |---------|----------| | Sprint status + findings | `sprints/SPRINT-003-...md` (updated) | | Taskitem-as-work-unit constraint | `docs/case_studies/whetstone_dsl_training_data_strategy.md` | | Intake contract constraint | `sprints/SPRINT-003-...md` (Architectural Constraint section) | | Split tool | `tools/split_prereq_op_dataset.py` | | TSV generator (RSA-sourced) | `tools/gen_prereq_op_rsa_data.py` | | 1000-run extracted dataset | `semantic/prereq_op_dataset_1000/` | | Trained baseline checkpoints | `/mnt/storage/fabricate_runs/whetstone_prereq_review_binary/` | | | `/mnt/storage/fabricate_runs/whetstone_prereq_approval_binary/` | ## What Is NOT Done (Next Session) **Execute SPRINT-004: Pipeline Decision Audit.** Full sprint document: `sprints/SPRINT-004-pipeline-decision-audit.md` The goal is a manual walkthrough of the full whetstone_DSL pipeline from raw project description to generated code. At every stage, record every decision being made, what inputs were available, and what the optimal method is (deterministic / RSA / SLM / LLM / human). Current LLM calls bundle many decisions together for token efficiency. The audit unbundles them. `generate_taskitems` alone contains 10+ separable decisions. Output: `docs/case_studies/whetstone_dsl_pipeline_decision_map.md` This document replaces the current gate inventory as the authoritative RSA input. Training data collection strategy follows from it, not before it. ### Start here next session Read `editor/src/mcp/RegisterArchitectIntakeTools.h` first — it is the pipeline entry point and its outputs constrain everything downstream. ## Do Not Do Next Session - Do not run any pipeline code - Do not write training data generators - Do not train any specialist - Do not modify whetstone_DSL source - Do not collect more run artifacts until the decision map is complete