Files
whetstone_RSA/presentations/rsa_tiny_models_talk.md

472 lines
7.8 KiB
Markdown
Raw Normal View History

# Tiny Models For Bounded NLP Decisions
`whetstone_RSA`
Fast-slide draft for a 5-10 minute programmer talk
This version is intentionally structured as many short slides with minimal text.
The goal is to support a visual talk with one idea per slide.
## Slide 1: Title
**Tiny Models For Bounded NLP Decisions**
`whetstone_RSA`
Turning messy user requests into deterministic tool calls
Visual idea:
Large title, small diagram of `user request -> contract -> tool`
Speaker notes:
This talk is about a middle layer between regex and LLMs.
## Slide 2: A Common Shape
Users type messy things.
Software needs clean actions.
Visual idea:
Messy speech bubble on left, clean typed API object on right
Speaker notes:
This shows up all over software.
## Slide 3: The Gap
We often have:
- fuzzy input
- bounded output
- deterministic execution
Visual idea:
Three stacked boxes with arrows
Speaker notes:
That middle combination is the important one.
## Slide 4: Todays Typical Options
- regex
- rules
- LLM prompt
Visual idea:
Three boxes, with regex looking brittle and LLM looking oversized
Speaker notes:
Most teams bounce between brittle and oversized.
## Slide 5: Regex Is Cheap
Regex and rules are:
- fast
- local
- inspectable
Visual idea:
Green checkmarks next to those three traits
Speaker notes:
This is why people keep reaching for rules.
## Slide 6: Regex Is Also Annoying
Regex and rules are:
- brittle
- tedious
- bad at paraphrases
Visual idea:
Same rule box cracking under many user phrasings
Speaker notes:
The moment users say the same thing many ways, rules get ugly.
## Slide 7: LLMs Feel Great
LLMs are:
- flexible
- natural-language friendly
- good with paraphrases
Visual idea:
Many different speech bubbles converging successfully
Speaker notes:
This is why people like them.
## Slide 8: LLMs Are Often Overkill
For bounded tasks, LLMs can be:
- expensive
- heavy
- hard to control
Visual idea:
Huge engine powering a tiny gear
Speaker notes:
A lot of current usage is more model than the task actually needs.
## Slide 9: The Missing Middle
There should be a middle layer.
Visual idea:
`regex -> RSA -> LLM`
Speaker notes:
That middle layer is what RSA is for.
## Slide 10: What RSA Is
RSA is a contract layer for bounded semantic decisions.
Visual idea:
Contract sheet icon between user and tool
Speaker notes:
Not a chatbot. Not a search engine. A decision contract layer.
## Slide 11: What RSA Does
RSA takes:
- messy input
- bounded schema
- optional context
Visual idea:
Three inputs feeding one box
Speaker notes:
The schema is what keeps it narrow.
## Slide 12: What RSA Returns
RSA returns:
- action
- slots
- confidence
- abstain or escalate
Visual idea:
Typed JSON-like result
Speaker notes:
The output is structured, not freeform prose.
## Slide 13: The Core Pattern
Natural language
-> semantic interpretation
-> bounded contract
-> deterministic software
Visual idea:
Simple four-step pipeline
Speaker notes:
The model interprets. The software executes.
## Slide 14: Example Query
`show me the failed builds from yesterday`
Visual idea:
One big centered query bubble
Speaker notes:
This is the kind of input users naturally want to type.
## Slide 15: Example Contract
Becomes:
- action: `search_builds`
- status: `failed`
- date: `yesterday`
- confidence: `0.93`
Visual idea:
Search query transformed into structured object
Speaker notes:
This is not generation. It is semantic interpretation into a contract.
## Slide 16: Example Execution
Then ordinary software:
- runs the query
- renders the results
- logs the action
Visual idea:
Database, search, and UI boxes after the contract
Speaker notes:
Once the contract exists, it is just normal software again.
## Slide 17: Why This Can Stay Tiny
The output is bounded.
Visual idea:
Huge input cloud mapped into a small action menu
Speaker notes:
The model does not need to generate arbitrary text.
## Slide 18: First Real Target
`whetstone_DSL`
AST-first codegen pipeline
Visual idea:
WhetstoneDSL box feeding RSA gates
Speaker notes:
This is not a toy example. These are real pipeline decisions.
## Slide 19: The Initial Gate
`prereq_op_selector`
Decides prerequisite operations
before a taskitem executes.
Visual idea:
Taskitem -> prereq gate -> operation checklist
Speaker notes:
The gate decides what must happen before execution.
## Slide 20: Gate Input
Input is structured task context:
- title
- requirements
- acceptance criteria
- constraints
Visual idea:
Four input cards into one gate
Speaker notes:
The model sees the compact decision context, not the whole world.
## Slide 21: Gate Output
Output was:
- `needs_resolve_dependencies`
- `needs_architect_review`
- `needs_validate_intake`
Visual idea:
Three boolean switches
Speaker notes:
This is the bounded contract surface.
## Slide 22: Source Corpus
1000 pipeline runs
JSON artifacts
real decision logs
Visual idea:
Stack of JSON files
Speaker notes:
The point is extracting supervision from actual pipeline behavior.
## Slide 23: Extraction
`extract_gate_rows.py`
turns run artifacts
into per-gate rows.
Visual idea:
JSON stack -> extraction script -> row table
Speaker notes:
Each row isolates one decision.
## Slide 24: Clean Rows
1321 total rows
24 schema-drift rows rejected
1297 clean rows
Visual idea:
Row counter with a small rejected pile
Speaker notes:
Bad rows get rejected instead of quietly poisoning the gate.
## Slide 25: The Surprise
`needs_validate_intake`
was always true.
1297 / 1297
Visual idea:
One switch permanently locked on
Speaker notes:
A label with zero variance is not learnable. It is a constant.
## Slide 26: Gate Decomposition
The real gate became two gates:
- resolve dependencies
- architect review
Visual idea:
One big gate splitting into two smaller gates
Speaker notes:
The data changed the gate shape. That is the process working.
## Slide 27: Tiny Specialists
One specialist per gate.
About 213K parameters.
About 800KB on disk.
Visual idea:
Tiny model file next to a normal image file
Speaker notes:
The size is small because the decision surface is small.
## Slide 28: First Results
Baseline accuracy:
- resolve: 69.4%
- architect: 75.0%
Visual idea:
Two simple bars
Speaker notes:
These are not victory-lap numbers. They are diagnostic numbers.
## Slide 29: Honest Diagnosis
1297 rows
but only about 17 unique text patterns.
Visual idea:
Lots of rows collapsing into a few templates
Speaker notes:
More rows from the same shape do not create richer signal.
## Slide 30: The Fix
Better input fields:
- architecture surface
- cross-component deps
- security sensitivity
- uncertainty score
Visual idea:
Taskitem schema before and after
Speaker notes:
The fix is improving the contract upstream, not pretending the weak model is good.
## Slide 31: Model Sizing
Match model size
to gate entropy.
Visual idea:
Gate difficulty -> model tier
Speaker notes:
The goal is the smallest model that clears the quality bar.
## Slide 32: What Makes A Gate Hard?
- more labels
- class imbalance
- wider context
- fuzzy boundaries
- slot interactions
Visual idea:
Difficulty knobs
Speaker notes:
These are the factors that should predict model tier.
## Slide 33: Runtime Scaffold
C++ library:
- gate definitions
- evidence
- diagnosis
- probes
Visual idea:
Fabricate specialist -> RSA runtime -> WhetstoneDSL
Speaker notes:
The repo is not just notes. There is a runtime scaffold.
## Slide 34: Careful Claim
RSA does not replace all LLM reasoning.
RSA replaces bounded semantic decisions.
Visual idea:
Bounded decisions inside scope, open-ended reasoning outside
Speaker notes:
Keep the claim narrow. That makes it harder to knock down.
## Slide 35: Closing
Fuzzy input in.
Typed contract out.
Deterministic software after that.
Visual idea:
Final clean pipeline graphic
Speaker notes:
Do not make the model own the whole task when it only needs to own the ambiguous boundary.