Five standalone HTML visuals covering the RSA pipeline, training pipeline, gate decomposition, dataset anatomy, and model tier sizing. Slide deck brief describing a 35-slide two-act deck for a data science audience. Handoff note documenting what was created this session. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
11 KiB
Slide Deck Brief — whetstone_RSA: Tiny Models for Bounded Decisions
Audience: Whetstone_RSA, ex data science instructor. Technical depth expected. Wants to understand the actual transformer training approach, not just the concept.
Format: Slide deck, dark theme preferred. One idea per slide, minimal text. Visuals are in the same directory as this file (HTML files — embed or screenshot them).
Length: ~30–35 slides across two acts.
Act I — The Problem and the Pattern (Slides 1–17)
These slides are already written in rsa_tiny_models_talk.md. Keep them as-is.
They establish the motivation: fuzzy input → bounded output → deterministic execution.
Key points already covered:
- Why regex is brittle
- Why LLMs are oversized for bounded tasks
- What RSA does (contract layer between user and tool)
- The core pipeline: input → RSA gate → typed contract → deterministic software
- Where this pattern shows up (command palettes, search bars, workflow routing)
Visual for Act I: visual_rsa_pipeline.html
Act II — The Actual Training Work (Slides 18–35)
This is the new material. It covers the real implementation, dataset, and results.
Slide 18: The Case Study
Title: First Real Target — whetstone_DSL
Content:
- whetstone_DSL is a C++ IDE with an AST-first codegen pipeline
- It uses an LLM to make many bounded decisions bundled together
- Goal: replace those bundles with one tiny specialist per decision gate
- First gate selected:
prereq_op_selector- Decides which prerequisite operations a taskitem needs before it can execute
Speaker notes: This is not a toy example. These gates run on real pipeline decisions.
Slide 19: What a Gate Is
Title: The Gate Contract
Content:
input: task title · requirements · acceptance criteria · constraints
output: needs_resolve_dependencies: bool
needs_architect_review: bool
needs_validate_intake: bool (always true — dropped)
confidence: float
abstain: bool
Visual idea: Simple box diagram: structured text in, typed struct out.
Speaker notes: The output is not generated text. It is a typed struct with known fields.
Slide 20: The Training Pipeline Overview
Title: From Run Artifacts to Specialist
Visual: visual_training_pipeline.html — use full slide.
Speaker notes: Six-step pipeline. Real data from real pipeline runs. The training framework (Fabricate) lives on the GPU desktop — can demo it live there.
Slide 21: Step 1 — The Corpus
Title: Source Data: 1000 Pipeline Runs
Content:
- whetstone_DSL was run 1000 times on real and synthetic projects
- Each run produces JSON log artifacts recording every taskitem decision
extract_gate_rows.pyparses these into per-gate supervision rows
Visual idea: Stack of JSON files on left → extraction tool → rows on right.
Slide 22: Step 2 — Gate Extraction
Title: Isolating One Decision at a Time
Content:
- Each row contains: decision_context · decision_label · decision_provenance
- Only the information available at decision time is allowed in the context
- Rows with schema-drift labels are rejected (24 of 1321 total)
- Result: 1297 clean rows for prereq_op_selector
Key extraction rule: Never leak downstream fields into the context. A gate must only see what was available when it would actually fire.
Slide 23: The Discovery — Gate Decomposition
Title: needs_validate_intake Is Always True
Content:
- Analysis of 1297 rows: needs_validate_intake = True in all 1297 of them
- A label with zero variance cannot be learned — it is a constant, not a gate
- Real gate decomposes into two independent binary classifiers:
- needs_resolve_dependencies (76.9% positive rate)
- needs_architect_review (19.8% positive rate)
Visual: visual_gate_decomposition.html
Speaker notes: This is a real finding, not a design choice made upfront. The data told us the gate shape was wrong. That's the process working correctly.
Slide 24: Dataset Anatomy
Title: What 1297 Rows Actually Looks Like
Visual: visual_dataset_anatomy.html — use full slide.
Speaker notes: The ~17 unique templates is the important number. High row count does not mean high diversity. This is the honest accounting.
Slide 25: The TSV Format
Title: Fabricate Training Format
Content:
label <TAB> hop <TAB> text
- label: 0 (no) or 1 (yes) — one file per binary gate
- hop: always 0 (unused routing field in Fabricate)
- text: composed from title · reasons · acceptance criteria · constraints
Four files produced:
- prereq_resolve_train.tsv (1102 rows)
- prereq_resolve_eval.tsv (195 rows)
- prereq_architect_train.tsv (1102 rows)
- prereq_architect_eval.tsv (195 rows)
Split: 85% train / 15% eval, random seed 42.
Slide 26: The Model — Tiny Transformer
Title: ~213K Parameters. ~800KB on Disk.
Content:
- Framework: Fabricate (custom tiny transformer training tool)
- Architecture: transformer encoder with binary classification head
- ~213K parameters per specialist
- Serialized checkpoint: ~800KB
- Training: GPU desktop, 4000 steps per specialist
- Inference: CPU-friendly, local, no network dependency
Visual idea: Size comparison — model checkpoint vs. a JPEG photo (~800KB each).
Speaker notes: This is the whole point. A model smaller than an average image file, making a bounded yes/no decision. Can demo inference live on the GPU desktop.
Slide 27: Training Results
Title: Baseline Results @ 4000 Steps
Content:
| Specialist | Accuracy | Steps | Positive rate |
|---|---|---|---|
| prereq_resolve (needs_resolve_dependencies) | 69.4% | 4000 | 76.9% |
| prereq_architect (needs_architect_review) | 75.0% | 4000 | 19.8% |
Important context:
- A majority classifier on needs_resolve would score ~77% (model is below naive baseline)
- A majority classifier on needs_architect would score ~80% (model is learning minority class)
- These are weak baselines — expected given the 17-template corpus
Visual idea: Bar chart with majority-classifier baseline marked as a reference line.
Slide 28: Why the Baselines Are Weak (And Why That's Fine)
Title: The Corpus Was Not Designed for This
Content:
- The taskitem schema was designed to record execution outputs, not discriminative inputs
- 1297 rows, but only ~17 unique text patterns — the model has almost no signal variation
- The labels are correct; the features are not rich enough yet
The fix is not more data from the same corpus. The fix is adding RSA-native fields to the taskitem schema upstream.
Speaker notes: This is an architectural finding. The pipeline methodology is validated. Signal quality is a schema problem, not a model problem.
Slide 29: What Would Make the Signal Better
Title: Better Input = Better Gate
Content: Each taskitem needs to carry:
architectural_surface— what part of the system is touchedcross_component_deps— explicit dependency flagssecurity_sensitive— flag for security-relevant changesuncertainty_score— explicit uncertainty signal
With these fields, some gates that currently need ML may become deterministic rules. The RSA entropy threshold shifts as the contract gets richer.
Visual idea: Taskitem schema before vs. after adding RSA fields.
Slide 30: The Entropy → Model Size Hypothesis
Title: Match the Model to the Decision Surface
Visual: visual_model_tiers.html
Speaker notes: The hypothesis is measurable and falsifiable. The experiment is: train each tier on the same gate, record the smallest one that clears the target.
Slide 31: Gate Entropy Factors
Title: What Makes a Gate Hard?
Content: Factors that increase required model tier:
- output_cardinality (more labels = harder)
- class_imbalance
- lexical_variation in inputs
- context_width_required
- label_boundary_fuzziness
- world_knowledge_dependence
- slot_interdependence
Target deliverable:
gate_id | entropy_score | smallest_passing_tier | accuracy | latency_ms
Speaker notes: We don't have this table yet. It's the next experiment.
Slide 32: What's Next — SPRINT-004
Title: Pipeline Decision Audit
Content: Before training more gates, map every decision in the full whetstone_DSL pipeline:
- Start at RegisterArchitectIntakeTools.h (pipeline entry point)
- At every stage: what decision is being made?
- What inputs are available?
- What is the optimal method: deterministic / RSA / SLM / LLM / human?
generate_taskitems alone bundles 10+ separable decisions.
Output: docs/case_studies/whetstone_dsl_pipeline_decision_map.md
Speaker notes: Train data collection strategy follows from the decision map. Don't collect data for the wrong gate shape.
Slide 33: What Is Already Deterministic
Title: Most of Whetstone Is Already Right
Content: Once the AST and target language are fixed, most current code generation is deterministic:
- Target generator dispatch — deterministic
- AST concept dispatch — deterministic
- Environment gating — deterministic
- C++ raising profile validation — deterministic
RSA is highest value in:
- taskitem generation
- execution contract shaping
- routing decisions
- pre-codegen profile decisions
Speaker notes: This is a feature, not a gap. Don't replace deterministic software with models. Find the actual ambiguity boundary.
Slide 34: The C++ Runtime
Title: whetstone_RSA Runtime Scaffold
Content: A C++ library wrapping the gate contracts:
GateDefinition·GateEvidence·GateDiagnosis·GateSuitabilityProbeRegistry+DiagnosisEngine- First concrete probe:
confidence_threshold_probe - Builds with CMake, demo executable runs
This is the deployment layer — gates are loaded here at runtime.
Visual idea: Stack diagram: Fabricate specialist → RSA runtime → whetstone_DSL pipeline.
Slide 35: The Summary
Title: What This Is
Content:
- Real, running pipeline decisions in whetstone_DSL
- Real extracted training data (1297 rows, two binary gates)
- Real trained specialists (~800KB each, ~213K params)
- Honest baselines (weak corpus, clear reason, clear fix)
- A measurable hypothesis: entropy → tier selection
- A C++ runtime that loads and runs them
Not a chatbot. Not a search engine. A bounded decision contract layer, with real data and real results.
Appendix Slides (optional, for Q&A)
- Rejection taxonomy (schema drift, sibling-uniform labels)
- Full gate inventory for whetstone_DSL
- The decision_contract_ontology (input_structure, output_topology, policy_stability, etc.)
- Backend selection taxonomy (transformer vs. gradient boosting vs. Bayes — kept separate from gate identity)
- The abstain/escalate policy design
Files in this directory
| File | Description |
|---|---|
rsa_tiny_models_talk.md |
Act I source (31 slides, already complete) |
slide_deck_brief.md |
This file |
visual_rsa_pipeline.html |
The core RSA pattern — fuzzy in, contract out |
visual_training_pipeline.html |
Six-step pipeline from run artifacts to trained specialist |
visual_gate_decomposition.html |
prereq_op_selector decomposed into two binary gates |
visual_dataset_anatomy.html |
Dataset stats, class balance, TSV format sample |
visual_model_tiers.html |
Entropy → model tier sizing diagram |