Files
whetstone_RSA/sprints/SPRINT-003-whetstone-dsl-gate-semantic-bootstrap.md
Bill Holcombe c4569de9ca Sprint 003 findings + Sprint 004 pipeline decision audit plan
- split_prereq_op_dataset.py: extract accepted/rejected/schema-drift rows per gate
- gen_prereq_op_rsa_data.py: convert accepted rows to Fabricate TSV format
- 1000-run extracted dataset: 1297 accepted prereq_op rows, 24 schema-drift
- Key finding: needs_validate_intake always True; effective gate is binary
  (needs_resolve_dependencies 77%, needs_architect_review 20%)
- Key finding: run corpus text is template-driven (17 unique inputs) — not
  suitable for text-based training; RSA exists before ideal input is available
- ROADMAP: gate lifecycle framing — goal is choice of backend, not forced
  determinism; RSA operates on imperfect/mixed/incomplete input by design
- training_data_strategy.md: taskitem-as-self-contained-work-unit constraint
- SPRINT-003: updated with 1000-run results, text audit findings, design constraints
- SPRINT-004: pipeline decision audit — manual walkthrough of every decision
  from project description to generated artifact; execute next session

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-12 21:43:17 -06:00

574 lines
17 KiB
Markdown

# Sprint 003: WhetstoneDSL Gate Semantic Bootstrap
## Status
Active — pending intake contract design (see Architectural Constraint below)
## Intent
Turn the current RSA architecture and gate inventory into an executable semantic
bootstrap focused on the first concrete `whetstone_DSL` case-study gates.
This sprint exists to close the gap between:
- generic RSA architecture
- generic semantic bootstrap language
- actual `whetstone_DSL` decision surfaces
The output of this sprint should be concrete enough that later work can proceed
without re-deriving what the first real gates are or how they map into semantic
entities.
## Architectural Constraint — Intake Contract
**Discovered this session.** This constraint changes how RSA training data is sourced.
The current whetstone_DSL run corpus predates a planned intake contract. The intended
design is that external tools (HiveMind, gleaner, constcad, etc.) submit project
descriptions in a **structured contract format** that explicitly carries the fields
needed to make RSA gate decisions bounded and classifiable.
Without the contract, the intake produces ~18 unique template outputs regardless of
project. With the contract, every project submission becomes a varied, labeled training
example with known decision-relevant fields.
**Correct pipeline (with contract):**
```
External tool
↓ submits structured project description (contract fields)
Intake contract schema ← explicit per-field semantics
architect_intake reads contract, produces normalized form
RSA gates fire on contract fields ← bounded, varied, discriminative
generate_taskitems ← deterministic given gate outputs
```
**Implications for RSA design:**
- Gate input features are **contract fields**, not free-form text
- Training data comes from **contract submissions**, not ad hoc specs
- The current corpus is valid for label schema validation only
- Sprint plan step descriptions remain a valid interim text source for baseline training
- The contract schema must be defined before RSA training data can be collected at scale
**Next design milestone (not this sprint):**
Define the intake contract schema: what fields must an external tool provide, and which
fields are the discriminative inputs for each immediate RSA gate (`verification_type`,
`worker_type`, `prereq_op_selector`, `cpp_raising_profile`).
## Why This Sprint Exists
`SPRINT-002-rsa-semantic-bootstrap.md` established the architectural direction,
but it is still too abstract to drive direct implementation.
We now have a concrete inventory in:
- `docs/case_studies/whetstone_dsl_gate_inventory.md`
The next step is to operationalize that inventory into:
- first concrete case-study gate records
- explicit semantic mapping rules
- clear enum-vs-structured-output boundaries
- acceptance criteria for what counts as a completed bootstrap
## Scope
This sprint is still semantic-first.
Primary outputs:
- a detailed sprint plan for the first case-study gate set
- semantic records for the first concrete `whetstone_DSL` gates
- mapping notes from `whetstone_DSL` decision surfaces into RSA entities
- explicit skeleton targets for later structured-output gates
This sprint does **not** require:
- model training
- wiring RSA runtime into `whetstone_DSL`
- replacing current `whetstone_DSL` heuristics
- projecting full Whetstone-authored AST entities yet
## Primary Gate Set
These are the first concrete RSA case-study gates for bootstrap:
### Immediate RSA gates
- `verification_type`
- `worker_type`
- `prereq_op_selector`
- `cpp_raising_profile`
### Skeleton-first semantic decisions
- `execution_contract_shape`
- `workstream_grouping`
- `target_file_selection`
## Gate Definitions To Lock Down
### 1. `verification_type`
Source surface:
- `RegisterArchitectIntakeTools.h`
Decision form:
- bounded enum
Expected outputs:
- `unit`
- `schema`
- `docs`
- `smoke`
Contract interpretation:
- unstructured textual input
- bounded discrete output
- stable policy
- deterministic downstream verification behavior
### 2. `worker_type`
Source surface:
- `TaskitemGeneratorV2.h`
Decision form:
- bounded enum
Expected outputs:
- `implementer`
- `reviewer`
- `architect`
- `qa`
Contract interpretation:
- unstructured or hybrid textual input
- bounded discrete output
- mixed context sensitivity
- deterministic downstream routing behavior
### 3. `prereq_op_selector`
Source surface:
- `TaskitemGeneratorV2.h`
Decision form:
- bounded structured output
- factorized / multi-label candidate
Expected atomic outputs:
- `validate-intake`
- `architect-review`
- `resolve-dependencies`
Contract interpretation:
- hybrid semantic gate
- not a good flat multiclass default
- primary probe target for factorization analysis
### 4. `cpp_raising_profile`
Source surface:
- `RegisterCppRaisingTools.h`
Decision form:
- bounded enum
Expected outputs:
- `safe-first`
- `perf-first`
- `interop-first`
Contract interpretation:
- structured symbolic input
- bounded discrete output
- deterministic downstream execution and artifact generation
## Skeleton Decisions To Represent But Not Train Yet
### `execution_contract_shape`
Why skeleton first:
- output is structured
- current semantics are implicit across many heuristics
- not yet stable enough as a direct training target
Required semantic fields:
- target files
- required tools
- acceptance commands
- step sequence
- cross-project targets
- resource locks
### `workstream_grouping`
Why skeleton first:
- current logic mixes lexical grouping with semantic bundling
- good future structured-output gate
- not yet clean enough for direct benchmark comparisons
### `target_file_selection`
Why skeleton first:
- current logic is bounded but still heavily heuristic
- needs clearer contract and evaluation rules before training
## Deliverables
### Deliverable A: Case-study gate records
Update the semantic bootstrap so it explicitly records:
- gate identity
- ontology fields
- output labels or slot schema
- candidate gate shape
- deterministic baseline availability
- downstream executor kind
for the four immediate gates.
Status:
- complete in `semantic/rsa_semantic_bootstrap_v0.json`
### Deliverable B: Mapping rules
Write down the mapping from current `whetstone_DSL` code surfaces into RSA
semantic entities:
- `TaskitemGeneratorV2`
- `TaskitemConfidenceAmbiguity`
- `RegisterArchitectIntakeTools`
- `RegisterCppRaisingTools`
- `RoutingEngine`
Status:
- complete in `docs/case_studies/whetstone_dsl_semantic_mapping.md`
### Deliverable C: Structured-output bootstrap boundary
Define how `prereq_op_selector` and `execution_contract_shape` are represented
without collapsing them into flat labels.
Status:
- partially complete
- `execution_contract_shape` skeleton captured in
`docs/case_studies/whetstone_dsl_execution_contract_shape_skeleton.md`
- extraction rules for `prereq_op_selector` and execution-contract child
decisions captured in
`docs/case_studies/whetstone_dsl_gate_extraction_specs_v0.md`
- machine-readable extraction contracts captured in
`semantic/rsa_gate_extraction_contracts_v0.json`
### Deliverable D: Artifact-grounded extraction path
Define which surviving `whetstone_DSL` corpora can support RSA-aligned
extraction, and which should remain context or evaluation corpora.
Status:
- partially complete
- audit captured in `docs/case_studies/whetstone_dsl_run_artifact_audit.md`
- first extraction rules captured in
`docs/case_studies/whetstone_dsl_gate_extraction_specs_v0.md`
- parser-facing extraction contracts captured in
`semantic/rsa_gate_extraction_contracts_v0.json`
## Active Next Task
Use `semantic/rsa_gate_extraction_contracts_v0.json` to draft the first
artifact parser for:
- `prereq_op_selector`
- `target_file_selection`
- `acceptance_command_selection`
- `required_tool_selection`
Status:
- complete
- parser implemented in `tools/extract_gate_rows.py`
- sample output generated in `semantic/extracted_gate_rows_sample.ndjson`
## Next Active Task
Run the parser over a broader taskitem-run slice and review extraction yield by:
- gate id
- provenance class
- rejection-flag frequency
Status:
- complete for first 50-run slice
- output captured in `semantic/extracted_gate_rows_slice.ndjson`
- yield review captured in
`docs/case_studies/whetstone_dsl_extraction_yield_snapshot.md`
## Updated Working Conclusion
The first extraction-ready dataset candidate is:
- `prereq_op_selector`
The following remain explicit but review-gated due to strong sibling-uniform
template behavior in the current slice:
- `target_file_selection`
- `acceptance_command_selection`
- `required_tool_selection`
## Next Active Task
Turn `prereq_op_selector` into the first accepted dataset slice, with:
- accepted rows
- rejected rows
- explicit schema-drift examples
Status:
- **complete** for the 50-run slice
- split produced by `tools/split_prereq_op_dataset.py`
- outputs at `semantic/prereq_op_dataset/`
- `accepted.ndjson` — 94 rows
- `rejected.ndjson` — 12 rows (all schema-drift)
- `schema_drift.ndjson` — 12 rows (same set; all rejections are schema-drift)
### Key Findings From Split
**Label distribution (accepted rows):**
| Combo | Count |
|-------|-------|
| validate\_intake only | 13 |
| validate\_intake + resolve\_dependencies | 47 |
| validate\_intake + architect\_review | 34 |
**`needs_validate_intake` is always True** across all 94 accepted rows.
This means:
- `validate_intake` is not a decision surface — it is a constant output in this corpus.
- The effective gate is a 3-way classification:
- `validate_only` (13 rows, 13.8%)
- `validate+resolve` (47 rows, 50%)
- `validate+architect` (34 rows, 36.2%)
- Or equivalently, two independent binary decisions:
- `needs_resolve_dependencies` (47/94 = 50%)
- `needs_architect_review` (34/94 = 36%)
**Schema-drift rows carry valid label signal.** The 12 rejected rows have correct
`needs_*` booleans. The only issue is `source_values` contains tool IDs mixed
with valid op names. Salvage path: re-extract only the valid op names and strip
the tool IDs, then accept the row.
### Decision Required Before Broadening
Before running over more runs, decide:
1. **Hard-filter or salvage schema-drift rows?**
- Hard-filter: discard 12 rows, keep extraction simple.
- Salvage: strip invalid values from `source_values`, promote to accepted.
- Recommendation: hard-filter for the first training run. Salvage later if
the accepted corpus stays small after broadening.
2. **Flat 3-way vs factorized binary formulation?**
- Flat 3-way: treat as single multiclass with 3 outputs (ignoring validate_intake).
- Factorized binary: two independent classifiers for `architect_review` and
`resolve_dependencies`. `validate_intake` is dropped as a constant.
- Recommendation: factorized binary. The two decisions appear semantically
orthogonal in the data.
### 1000-Run Extraction Results
Output at `semantic/prereq_op_dataset_1000/` (1000 runs, 1321 tasks):
- accepted: 1297 rows
- rejected: 24 rows (all schema-drift)
- schema drift variants:
- tool-ID drift (`whetstone_generate_taskitems`, etc.): 12 rows
- `manual-approval` drift: 12 rows (undocumented op name in `prerequisiteOps`)
**Label distribution (1000-run accepted):**
| Combo | Count | Pct |
|-------|-------|-----|
| validate\_intake + resolve\_dependencies | 996 | 76.8% |
| validate\_intake + architect\_review | 256 | 19.7% |
| validate\_intake only | 44 | 3.4% |
| all three | 1 | 0.1% |
`needs_validate_intake` remains always True at 1000-run scale.
Effective gate at 1000 runs:
- `needs_resolve_dependencies`: 997/1297 = 76.9% positive
- `needs_architect_review`: 257/1297 = 19.8% positive
Class imbalance for `needs_architect_review` (~80/20) is manageable for training.
The `validate_intake` constant can be dropped from the label space.
Both schema-drift variants are **salvageable**: valid bits (`needs_*`) are correct;
only `source_values` contains noise. Recommendation: hard-filter for first training
run, evaluate whether 24 extra rows are worth a salvage pass after.
### Text Feature Audit — Critical Finding
Attempted text composition from extracted row context fields: `title`, `reasons`,
`normalized_requirements`. Result:
- unique texts across 1297 rows: **17** (title+reasons), **18** (acceptance criteria)
- top text appears **697 times** ("Execution Readiness Primary | high_confidence_clear_path")
- structured features (confidence, dependency count) are also nearly constant
**Root cause:** The whetstone_DSL corpus uses a small set of canonical intake specs
(~18 unique templates) run many times. All tasks generated from the same intake share
the same context text. The step description text — the actual discriminative feature —
is not stored in run artifacts.
**Conclusion:**
The run corpus captures **labels** correctly but not the **text features** needed to
train a text-based prereq_op specialist.
The correct training data source for text input remains sprint plan step descriptions
(as used by `gen_prereq_op_data.py` in the specialist fleet). The run corpus contributes:
- validated label schema
- label frequency distribution (validate-intake constant, resolve 77%, architect 20%)
- schema-drift pattern documentation
- corpus for future label-correction of heuristic sprint-plan data
**Data strategy going forward:**
| Source | Role |
|--------|------|
| Sprint plan step descriptions | Primary text input for training |
| Run corpus (`accepted.ndjson`) | Label validation and frequency baseline |
| Run corpus schema-drift rows | Schema-drift taxonomy for robustness testing |
The sprint plan approach is already implemented in:
`CLionProjects/whetstone_DSL/specialists/scripts/gen_prereq_op_data.py`
That script uses keyword heuristics to assign labels. The RSA extraction work validates
that the heuristic label schema is correct; improving label accuracy requires linking
sprint plan steps back to run artifacts (not yet feasible).
### Next Active Task
Proceed to training using existing sprint plan data with the factorized binary formulation:
1. Run `gen_prereq_op_data.py` to regenerate sprint plan TSVs.
2. Split into two binary TSVs (resolve / architect) using `prereq_op_binary_split.py`.
3. Train both specialists with `train_specialist_pt.py`.
4. Evaluate combined accuracy against the 4-way eval set.
The RSA-extracted `accepted.ndjson` dataset remains available as a label-validation
reference and can be used for post-training audit (do sprint-plan-trained specialists
generalize to real run labels?).
- `execution_contract_shape` skeleton is now defined in
`docs/case_studies/whetstone_dsl_execution_contract_shape_skeleton.md`
- `target_file_selection` and `workstream_grouping` still need explicit skeleton field sets
### Deliverable D: Probe-readiness notes
Record which probes should apply first to each gate:
- `confidence_threshold_probe`
- `factorization_probe`
- `schema_stability_probe`
- `deterministic_rule_probe`
- `context_width_probe`
Status:
- partially complete
- immediate-gate probe candidates are recorded in the gate inventory and
bootstrap notes
- still needs a single consolidated probe-readiness artifact
## Recommended Execution Order
1. Extend the semantic bootstrap artifact with the first four case-study gates.
2. Record mapping notes from each current `whetstone_DSL` implementation surface.
3. Represent `prereq_op_selector` as structured/factorized rather than flat.
4. Represent skeleton-only decisions for execution-contract shaping and grouping.
5. Record which current gates should remain deterministic.
6. Leave runtime wiring and training for later sprints.
Current progress:
- steps 1, 2, and 3 are complete
- `execution_contract_shape` is complete enough for first-pass skeleton treatment
- `target_file_selection` is now the active task
## Acceptance Criteria
- the sprint names a fixed first gate set rather than speaking about gates in the abstract
- the semantic bootstrap artifact contains explicit records for:
- `verification_type`
- `worker_type`
- `prereq_op_selector`
- `cpp_raising_profile`
- each gate is classified as enum, structured, or factorized
- at least three skeleton-first decision entities are identified and justified
- deterministic-baseline gates are explicitly separated from RSA gates
- the sprint is actionable without reopening architecture debate
## Not In Scope
- training a new specialist
- modifying `whetstone_DSL` runtime code
- adding new RSA C++ runtime classes unless needed for semantic alignment
- making Whetstone AST the full live authoring source this sprint
## Exit Condition
This sprint is complete when a later session can answer all of the following
without ambiguity:
- what are the first real `whetstone_DSL` gates in RSA terms
- which are enum versus structured-output
- which should remain deterministic
- which structured decisions need skeleton treatment before training
- what the first semantic bootstrap records look like
## Immediate First Task
Extend `semantic/rsa_semantic_bootstrap_v0.json` with a case-study gate section
covering:
- `verification_type`
- `worker_type`
- `prereq_op_selector`
- `cpp_raising_profile`