7.8 KiB
Tiny Models For Bounded NLP Decisions
whetstone_RSA
Fast-slide draft for a 5-10 minute programmer talk
This version is intentionally structured as many short slides with minimal text. The goal is to support a visual talk with one idea per slide.
Slide 1: Title
Tiny Models For Bounded NLP Decisions
whetstone_RSA
Turning messy user requests into deterministic tool calls
Visual idea:
Large title, small diagram of user request -> contract -> tool
Speaker notes: This talk is about a middle layer between regex and LLMs.
Slide 2: A Common Shape
Users type messy things.
Software needs clean actions.
Visual idea: Messy speech bubble on left, clean typed API object on right
Speaker notes: This shows up all over software.
Slide 3: The Gap
We often have:
- fuzzy input
- bounded output
- deterministic execution
Visual idea: Three stacked boxes with arrows
Speaker notes: That middle combination is the important one.
Slide 4: Today’s Typical Options
- regex
- rules
- LLM prompt
Visual idea: Three boxes, with regex looking brittle and LLM looking oversized
Speaker notes: Most teams bounce between brittle and oversized.
Slide 5: Regex Is Cheap
Regex and rules are:
- fast
- local
- inspectable
Visual idea: Green checkmarks next to those three traits
Speaker notes: This is why people keep reaching for rules.
Slide 6: Regex Is Also Annoying
Regex and rules are:
- brittle
- tedious
- bad at paraphrases
Visual idea: Same rule box cracking under many user phrasings
Speaker notes: The moment users say the same thing many ways, rules get ugly.
Slide 7: LLMs Feel Great
LLMs are:
- flexible
- natural-language friendly
- good with paraphrases
Visual idea: Many different speech bubbles converging successfully
Speaker notes: This is why people like them.
Slide 8: LLMs Are Often Overkill
For bounded tasks, LLMs can be:
- expensive
- heavy
- hard to control
Visual idea: Huge engine powering a tiny gear
Speaker notes: A lot of current usage is more model than the task actually needs.
Slide 9: The Missing Middle
There should be a middle layer.
Visual idea:
regex -> RSA -> LLM
Speaker notes: That middle layer is what RSA is for.
Slide 10: What RSA Is
RSA is a contract layer for bounded semantic decisions.
Visual idea: Contract sheet icon between user and tool
Speaker notes: Not a chatbot. Not a search engine. A decision contract layer.
Slide 11: What RSA Does
RSA takes:
- messy input
- bounded schema
- optional context
Visual idea: Three inputs feeding one box
Speaker notes: The schema is what keeps it narrow.
Slide 12: What RSA Returns
RSA returns:
- action
- slots
- confidence
- abstain or escalate
Visual idea: Typed JSON-like result
Speaker notes: The output is structured, not freeform prose.
Slide 13: The Core Pattern
Natural language
-> semantic interpretation
-> bounded contract
-> deterministic software
Visual idea: Simple four-step pipeline
Speaker notes: The model interprets. The software executes.
Slide 14: Example Query
show me the failed builds from yesterday
Visual idea: One big centered query bubble
Speaker notes: This is the kind of input users naturally want to type.
Slide 15: Example Contract
Becomes:
- action:
search_builds - status:
failed - date:
yesterday - confidence:
0.93
Visual idea: Search query transformed into structured object
Speaker notes: This is not generation. It is semantic interpretation into a contract.
Slide 16: Example Execution
Then ordinary software:
- runs the query
- renders the results
- logs the action
Visual idea: Database, search, and UI boxes after the contract
Speaker notes: Once the contract exists, it is just normal software again.
Slide 17: Why This Can Stay Tiny
The output is bounded.
Visual idea: Huge input cloud mapped into a small action menu
Speaker notes: The model does not need to generate arbitrary text.
Slide 18: First Real Target
whetstone_DSL
AST-first codegen pipeline
Visual idea: WhetstoneDSL box feeding RSA gates
Speaker notes: This is not a toy example. These are real pipeline decisions.
Slide 19: The Initial Gate
prereq_op_selector
Decides prerequisite operations
before a taskitem executes.
Visual idea: Taskitem -> prereq gate -> operation checklist
Speaker notes: The gate decides what must happen before execution.
Slide 20: Gate Input
Input is structured task context:
- title
- requirements
- acceptance criteria
- constraints
Visual idea: Four input cards into one gate
Speaker notes: The model sees the compact decision context, not the whole world.
Slide 21: Gate Output
Output was:
needs_resolve_dependenciesneeds_architect_reviewneeds_validate_intake
Visual idea: Three boolean switches
Speaker notes: This is the bounded contract surface.
Slide 22: Source Corpus
1000 pipeline runs
JSON artifacts
real decision logs
Visual idea: Stack of JSON files
Speaker notes: The point is extracting supervision from actual pipeline behavior.
Slide 23: Extraction
extract_gate_rows.py
turns run artifacts
into per-gate rows.
Visual idea: JSON stack -> extraction script -> row table
Speaker notes: Each row isolates one decision.
Slide 24: Clean Rows
1321 total rows
24 schema-drift rows rejected
1297 clean rows
Visual idea: Row counter with a small rejected pile
Speaker notes: Bad rows get rejected instead of quietly poisoning the gate.
Slide 25: The Surprise
needs_validate_intake
was always true.
1297 / 1297
Visual idea: One switch permanently locked on
Speaker notes: A label with zero variance is not learnable. It is a constant.
Slide 26: Gate Decomposition
The real gate became two gates:
- resolve dependencies
- architect review
Visual idea: One big gate splitting into two smaller gates
Speaker notes: The data changed the gate shape. That is the process working.
Slide 27: Tiny Specialists
One specialist per gate.
About 213K parameters.
About 800KB on disk.
Visual idea: Tiny model file next to a normal image file
Speaker notes: The size is small because the decision surface is small.
Slide 28: First Results
Baseline accuracy:
- resolve: 69.4%
- architect: 75.0%
Visual idea: Two simple bars
Speaker notes: These are not victory-lap numbers. They are diagnostic numbers.
Slide 29: Honest Diagnosis
1297 rows
but only about 17 unique text patterns.
Visual idea: Lots of rows collapsing into a few templates
Speaker notes: More rows from the same shape do not create richer signal.
Slide 30: The Fix
Better input fields:
- architecture surface
- cross-component deps
- security sensitivity
- uncertainty score
Visual idea: Taskitem schema before and after
Speaker notes: The fix is improving the contract upstream, not pretending the weak model is good.
Slide 31: Model Sizing
Match model size
to gate entropy.
Visual idea: Gate difficulty -> model tier
Speaker notes: The goal is the smallest model that clears the quality bar.
Slide 32: What Makes A Gate Hard?
- more labels
- class imbalance
- wider context
- fuzzy boundaries
- slot interactions
Visual idea: Difficulty knobs
Speaker notes: These are the factors that should predict model tier.
Slide 33: Runtime Scaffold
C++ library:
- gate definitions
- evidence
- diagnosis
- probes
Visual idea: Fabricate specialist -> RSA runtime -> WhetstoneDSL
Speaker notes: The repo is not just notes. There is a runtime scaffold.
Slide 34: Careful Claim
RSA does not replace all LLM reasoning.
RSA replaces bounded semantic decisions.
Visual idea: Bounded decisions inside scope, open-ended reasoning outside
Speaker notes: Keep the claim narrow. That makes it harder to knock down.
Slide 35: Closing
Fuzzy input in.
Typed contract out.
Deterministic software after that.
Visual idea: Final clean pipeline graphic
Speaker notes: Do not make the model own the whole task when it only needs to own the ambiguous boundary.