254 lines
11 KiB
Markdown
254 lines
11 KiB
Markdown
|
|
# Phase 6: Library Dispatch — Sprint Plan
|
|||
|
|
**Steps 1958–1982 | Sprints 286–290**
|
|||
|
|
**Written:** 2026-03-02
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Goal
|
|||
|
|
|
|||
|
|
Transform the task execution pipeline from **vanilla language programming** into
|
|||
|
|
**library-aware, per-task deterministic dispatch**. After Phase 6, every taskitem
|
|||
|
|
execution contract specifies not just a language but the exact library and API
|
|||
|
|
function to use for that operation — chosen deterministically, with no LLM or
|
|||
|
|
human in the loop for that decision.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## The Problem Being Solved
|
|||
|
|
|
|||
|
|
Current pipeline output (after Phase 5):
|
|||
|
|
```
|
|||
|
|
taskitem: "implement JSON serialization for UserRecord"
|
|||
|
|
executionContract: { language: "C++", stepId: "serialize" }
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Target pipeline output (after Phase 6):
|
|||
|
|
```
|
|||
|
|
taskitem: "implement JSON serialization for UserRecord"
|
|||
|
|
executionContract: {
|
|||
|
|
language: "C++",
|
|||
|
|
library: "nlohmann_json",
|
|||
|
|
operationDomain: "serialization.json",
|
|||
|
|
preferredAPIs: ["nlohmann::json::dump()", "nlohmann::json::parse()"],
|
|||
|
|
capabilityScore: 0.92,
|
|||
|
|
avoidedLibraries: [{ name: "protobuf", reason: "no-pretty-print, slow-large-arrays" }],
|
|||
|
|
justification: "nlohmann scores 0.92 vs protobuf 0.60 for serialization.json"
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The SLM receives a pre-decided contract. Its only job: write syntax around a
|
|||
|
|
pre-specified API call. Smallest possible decision space → highest possible
|
|||
|
|
accuracy → closest to determinism.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Architecture
|
|||
|
|
|
|||
|
|
### New Components
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
OperationTaxonomy — fixed registry of operation domains
|
|||
|
|
anchored to what spec writers naturally write
|
|||
|
|
e.g. serialization.json, compute.matrix, network.http.client,
|
|||
|
|
io.stream.audio, crypto.hash, db.query.sql, ...
|
|||
|
|
|
|||
|
|
LibraryCapabilityRecord — per (library, operationDomain):
|
|||
|
|
{ score: float, preferredAPIs: [], knownWeaknesses: [],
|
|||
|
|
notApplicable: bool, annotatedAt: date, source: str }
|
|||
|
|
|
|||
|
|
LibraryCapabilityLedger — registry of LibraryCapabilityRecords
|
|||
|
|
keyed by (libraryId, operationDomain)
|
|||
|
|
one-time annotation, deterministic reads forever
|
|||
|
|
|
|||
|
|
PerTaskLibrarySelector — given (normalizedRequirements, availableLibraries):
|
|||
|
|
for each operation domain in the task:
|
|||
|
|
rank available libraries by score
|
|||
|
|
return LibrarySelection { library, domain, score,
|
|||
|
|
preferredAPIs, avoidedLibraries }
|
|||
|
|
fully deterministic: sort by score, pick max, log justification
|
|||
|
|
|
|||
|
|
LibrarySymbolAdvisor — given (library, operationDomain):
|
|||
|
|
return specific API symbols for that operation
|
|||
|
|
seeded from: LSP symbol index, pre-built catalog, or both
|
|||
|
|
deterministic lookup, no inference required
|
|||
|
|
|
|||
|
|
[Extended] generate_taskitems — execution contract gains:
|
|||
|
|
selectedLibrary, operationDomain, capabilityScore,
|
|||
|
|
preferredAPIs, avoidedLibraries, justification
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Selection Pipeline (fully deterministic after annotation)
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
spec text
|
|||
|
|
→ architect_intake normalize requirements, extract operation domains
|
|||
|
|
→ language fitness scorer pick language per operation (Phase 1, complete)
|
|||
|
|
→ library capability ledger pick library per operation domain
|
|||
|
|
(deterministic: max score lookup)
|
|||
|
|
→ library symbol advisor pick specific API per operation
|
|||
|
|
(deterministic: catalog lookup)
|
|||
|
|
→ generate_taskitems execution contract: language + library + APIs
|
|||
|
|
→ SLM / code generator fills in syntax only — no decisions required
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Operation Taxonomy Design Principles
|
|||
|
|
|
|||
|
|
1. **Anchored to spec language** — domains are what a spec writer writes, not what
|
|||
|
|
a library author writes. "serialize to JSON" not "invoke jackson.ObjectMapper".
|
|||
|
|
|
|||
|
|
2. **Domain granularity** — coarser than individual functions, finer than "IO":
|
|||
|
|
- `serialization.json` ≠ `serialization.binary` ≠ `serialization.xml`
|
|||
|
|
- `compute.matrix` ≠ `compute.ml.training` ≠ `compute.ml.inference`
|
|||
|
|
- `network.http.client` ≠ `network.http.server` ≠ `network.websocket`
|
|||
|
|
|
|||
|
|
3. **Hierarchical** — `serialization.json` is a child of `serialization`,
|
|||
|
|
enabling fallback scoring when a library annotates at the parent level.
|
|||
|
|
|
|||
|
|
4. **Finite and versioned** — the taxonomy is a versioned registry, not free text.
|
|||
|
|
New domains are added deliberately, not automatically.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Library Capability Scoring Principles
|
|||
|
|
|
|||
|
|
- **Score range:** 0.00 (not applicable) to 1.00 (reference implementation)
|
|||
|
|
- **Score meaning:**
|
|||
|
|
- 0.00 — not applicable for this domain
|
|||
|
|
- 0.60–0.75 — functional but with significant known weaknesses
|
|||
|
|
- 0.76–0.90 — good, minor weaknesses or API ergonomic issues
|
|||
|
|
- 0.91–1.00 — preferred reference for this domain
|
|||
|
|
- **Annotation sources** (in priority order):
|
|||
|
|
1. Formal benchmark data (throughput, latency, memory, correctness coverage)
|
|||
|
|
2. Known CVE / correctness gaps from public vulnerability databases
|
|||
|
|
3. API surface analysis (does a direct function exist for the operation?)
|
|||
|
|
4. Community consensus from spec sheets / documentation
|
|||
|
|
- **Annotation is one-time authoring** — the selector runs deterministically forever
|
|||
|
|
- **Weaknesses are first-class** — a library can score 0.95 for binary serialization
|
|||
|
|
and explicitly 0.60 for JSON serialization; both entries exist in the ledger
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Sprint Breakdown
|
|||
|
|
|
|||
|
|
### Sprint 286 — OperationTaxonomy + LibraryCapabilityLedger (steps 1958–1962)
|
|||
|
|
|
|||
|
|
| Step | Component | Description |
|
|||
|
|
|------|-----------|-------------|
|
|||
|
|
| 1958 | OperationDomain | Enum/registry of operation domains with hierarchy; lookup, parent traversal |
|
|||
|
|
| 1959 | LibraryCapabilityRecord | Struct + validator; score range enforcement, weakness list, API list |
|
|||
|
|
| 1960 | LibraryCapabilityLedger | Registry keyed by (libraryId, operationDomain); get/set/has/score |
|
|||
|
|
| 1961 | OperationTaxonomy | Full domain tree with 40+ initial entries; classify(text)→domain |
|
|||
|
|
| 1962 | Integration | Seed ledger with 8 real libraries × their operation domains; verify lookups |
|
|||
|
|
|
|||
|
|
**Initial library set for seeding:**
|
|||
|
|
- `nlohmann_json` — serialization.json: 0.92
|
|||
|
|
- `protobuf` — serialization.binary: 0.95, serialization.json: 0.60
|
|||
|
|
- `serde_json` (Rust) — serialization.json: 0.98
|
|||
|
|
- `cublas` — compute.matrix: 0.99 (GPU), compute.matrix: 0.00 (CPU fallback)
|
|||
|
|
- `thrust` — compute.parallel: 0.97
|
|||
|
|
- `numpy` — compute.matrix: 0.95, compute.statistics: 0.93
|
|||
|
|
- `tokio` (Rust) — network.async: 0.97, io.stream: 0.95
|
|||
|
|
- `boost.asio` — network.async: 0.91, network.http.client: 0.78
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
### Sprint 287 — PerTaskLibrarySelector (steps 1963–1967)
|
|||
|
|
|
|||
|
|
| Step | Component | Description |
|
|||
|
|
|------|-----------|-------------|
|
|||
|
|
| 1963 | LibrarySelection | Struct: library, domain, score, preferredAPIs, avoidedLibraries, justification |
|
|||
|
|
| 1964 | OperationClassifier | Classifies normalized requirement text → OperationDomain (deterministic) |
|
|||
|
|
| 1965 | PerTaskLibrarySelector | Core selector: rank by score, return best + avoided list + justification |
|
|||
|
|
| 1966 | SelectionJustificationLog | Structured log of every selection decision (auditable, deterministic replay) |
|
|||
|
|
| 1967 | Integration | 8-task spec with mixed operations; verify correct library chosen per task |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
### Sprint 288 — LibrarySymbolAdvisor (steps 1968–1972)
|
|||
|
|
|
|||
|
|
| Step | Component | Description |
|
|||
|
|
|------|-----------|-------------|
|
|||
|
|
| 1968 | LibrarySymbolRecord | Struct: functionName, signature, operationDomain, usagePattern, caveats |
|
|||
|
|
| 1969 | LibrarySymbolCatalog | Registry of LibrarySymbolRecords; lookup by (library, operationDomain) |
|
|||
|
|
| 1970 | LSPSymbolExtractor | Parses LSP completion/hover responses → LibrarySymbolRecords |
|
|||
|
|
| 1971 | LibrarySymbolAdvisor | Given (library, operationDomain) → ranked symbol recommendations |
|
|||
|
|
| 1972 | Integration | CUDA: advisee returns cublasSgemm for compute.matrix; nlohmann: json::parse |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
### Sprint 289 — Integration into generate_taskitems (steps 1973–1977)
|
|||
|
|
|
|||
|
|
| Step | Component | Description |
|
|||
|
|
|------|-----------|-------------|
|
|||
|
|
| 1973 | EnrichedExecutionContract | Extends existing contract with library dispatch fields |
|
|||
|
|
| 1974 | TaskitemLibraryAnnotator | Runs selector + advisor per taskitem; injects into contract |
|
|||
|
|
| 1975 | AnnotatedTaskitemValidator | Validates library fields present when domain is known |
|
|||
|
|
| 1976 | DispatchJustificationReport | Per-project report: every task → library choice → score → why |
|
|||
|
|
| 1977 | Integration | Full pipeline: spec → intake → selector → advisor → enriched taskitems |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
### Sprint 290 — CUDA End-to-End Proof (steps 1978–1982)
|
|||
|
|
|
|||
|
|
| Step | Component | Description |
|
|||
|
|
|------|-----------|-------------|
|
|||
|
|
| 1978 | CUDALibraryProfile | Full CUDA ledger entries: cublas, cufft, thrust, cudnn domains + scores |
|
|||
|
|
| 1979 | CUDASymbolCatalog | 20+ CUDA API symbols with signatures and usage patterns |
|
|||
|
|
| 1980 | ProjectStackDeclarator | Declares available libraries for a project (input to selector) |
|
|||
|
|
| 1981 | ProofPipelineRunner | Runs full Phase 6 pipeline on a CUDA project spec |
|
|||
|
|
| 1982 | Integration | Spec: "GPU matrix multiply + JSON result serialization"
|
|||
|
|
Output: cublasSgemm for compute.matrix, nlohmann for serialization.json
|
|||
|
|
Sprint290IntegrationSummary |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## What Changes for the SLM After Phase 6
|
|||
|
|
|
|||
|
|
**Before Phase 6** — SLM decides:
|
|||
|
|
- What library to use (reasoning required)
|
|||
|
|
- What API to call (knowledge required)
|
|||
|
|
- Which edge cases the library handles (judgment required)
|
|||
|
|
|
|||
|
|
**After Phase 6** — SLM receives:
|
|||
|
|
```
|
|||
|
|
language: C++
|
|||
|
|
library: cublas
|
|||
|
|
operationDomain: compute.matrix
|
|||
|
|
preferredAPIs: ["cublasSgemm(handle, CUBLAS_OP_N, CUBLAS_OP_N, m, n, k, &alpha, A, lda, B, ldb, &beta, C, ldc)"]
|
|||
|
|
knownWeaknesses: ["requires-device-memory-management", "no-automatic-handle-lifecycle"]
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
SLM job: write the function body that calls this API with these parameters.
|
|||
|
|
No reasoning. No library knowledge. Syntax only.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Impact on Training Data
|
|||
|
|
|
|||
|
|
After Phase 6 is live:
|
|||
|
|
- **Retire:** `mcp_calls_good.jsonl`, `mcp_call_quality.jsonl` — these teach
|
|||
|
|
the SLM to make decisions that are now deterministic. Training on them teaches
|
|||
|
|
wrong behavior.
|
|||
|
|
- **Archive:** `mcp_calls_bad.jsonl` — failure modes are architecture-independent.
|
|||
|
|
- **Keep:** `taskitem_pipeline_runs.jsonl` — pipeline structure signal remains valid.
|
|||
|
|
- **Keep:** `datasets/` benchmark specs — reusable eval sets regardless of architecture.
|
|||
|
|
|
|||
|
|
New training data target post-Phase 6: syntax-fill examples where the contract
|
|||
|
|
is fully specified and the SLM is evaluated only on correctness of the generated
|
|||
|
|
function body — not on the library/API decision.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Success Criteria for Phase 6
|
|||
|
|
|
|||
|
|
1. Given a project spec with a declared library stack, every taskitem execution
|
|||
|
|
contract contains: `selectedLibrary`, `operationDomain`, `capabilityScore`,
|
|||
|
|
`preferredAPIs`, `avoidedLibraries`, `justification`
|
|||
|
|
2. Selection is deterministic: same spec + same ledger = same output, always
|
|||
|
|
3. No LLM or human consulted during selection
|
|||
|
|
4. CUDA project proof: "GPU matrix multiply" → `cublasSgemm`, not a kernel loop
|
|||
|
|
5. Mixed-stack proof: project uses protobuf for IPC but `serialization.json` task
|
|||
|
|
correctly selects `nlohmann_json` over protobuf (0.92 vs 0.60)
|