# WhetstoneDSL Training Data Strategy ## Purpose This document records the data-handling rule for turning existing `whetstone_DSL` outputs into RSA training data. The immediate concern is that much of the existing material may be: - whole JSON responses - whole taskitem packets - multi-decision LLM outputs rather than clean one-decision-per-example supervision. That means the data cannot be used naively. ## Core Rule Do not train per-gate specialists directly on full taskitem or full-response outputs without reconstructing: - the exact local context for the decision - the specific decision target being learned - whether the label was independently chosen or entangled with other choices ## Why This Matters The old pipeline was optimized for: - one LLM or SLM making several bounded choices together - minimizing token use by sharing context across those choices The new RSA direction is different: - one specialist per bounded decision - narrower context per decision - explicit gate contracts So existing outputs may reflect: - coupled decisions - prompt-order effects - hidden shared rationale - labels that were only valid because other fields in the same response were chosen jointly If we ignore that, the training data will be semantically wrong even if it looks large. ## Safe Extraction Rule For each candidate gate, build examples as: - `decision_context` - `decision_label` - `decision_provenance` and not as: - full response blob -> label ### `decision_context` Must include only the information that would have been available at the moment that gate should fire. Examples: - `worker_type` - task title - task intent - maybe requirement summary - not downstream fields derived after routing - `verification_type` - target files - task title - task intent - not later acceptance-command decisions unless those are part of the gate by contract - `prereq_op_selector` - task description - uncertainty and dependency signals if the intended gate contract includes them - not later execution outcomes ### `decision_label` Must match the new gate contract exactly. If the historic data uses a different shape than the new gate, do not force it. Instead: - transform it carefully - or discard it - or redefine the gate ### `decision_provenance` Each extracted example should track: - source artifact - source run id if available - whether label was explicit or inferred - whether the example came from: - human-authored heuristic baseline - LLM output - SLM output - deterministic reconstruction This is necessary because not all labels have equal trust. ## Gate-Specific Warnings ### `worker_type` This is likely recoverable from older taskitems, but only if the worker role was explicitly present and chosen before downstream execution shaping. Risk: - label polluted by later execution-contract logic ### `verification_type` This is relatively safe because it can often be reconstructed from target files and acceptance intent. Risk: - historic labels may include joint decisions with acceptance-command shaping ### `prereq_op_selector` This is the riskiest early gate for naive extraction. Reason: - older systems may have emitted a flat list of ops as one bundle - the new semantic contract treats this as structured or factorized Rule: - do not flatten structured op bundles into arbitrary multiclass labels ### `cpp_raising_profile` This is likely only useful if old runs actually recorded explicit profile choice with enough IR context. If not, this gate may need to start with synthetic or policy-authored data rather than mined response logs. ## Easy/Challenging Project Corpora The `100 easy` and `25 challenging` project sets are valuable, but they are not automatically specialist training sets. They are best treated as: - context reservoirs - replay/evaluation corpora - example generators for extracting many per-gate decisions The right pattern is: 1. replay or parse the project artifact 2. isolate one decision surface at a time 3. extract only the local context relevant to that gate 4. record the corresponding bounded label This should produce a gate-native dataset rather than a project-native dataset. ## Compatibility Constraint We should not assume the new specialists must fit perfectly into the exact old taskitem bundle shape. There are two valid outcomes: ### Option A: Fit specialists into current taskitems Use when: - the old taskitem structure already matches clean gate boundaries - the decision context can be isolated without distortion ### Option B: Rebuild taskitems around explicit gates Use when: - the current taskitems bundle too many coupled decisions - the old layout hides which context belongs to which choice - structured gates are being forced into flat labels This is not failure. It may be the correct result of moving from one-model-many-decisions to many-model-one-decision architecture. ## Recommended Next Step Before mining historic runs, define extraction specs for the first four gates: - `verification_type` - `worker_type` - `prereq_op_selector` - `cpp_raising_profile` Each extraction spec should state: - allowed input fields - forbidden leaked fields - label shape - trust level of recovered labels Only after that should we parse old JSON outputs or taskitem runs.