254 lines
11 KiB
Markdown
254 lines
11 KiB
Markdown
# Phase 6: Library Dispatch — Sprint Plan
|
||
**Steps 1958–1982 | Sprints 286–290**
|
||
**Written:** 2026-03-02
|
||
|
||
---
|
||
|
||
## Goal
|
||
|
||
Transform the task execution pipeline from **vanilla language programming** into
|
||
**library-aware, per-task deterministic dispatch**. After Phase 6, every taskitem
|
||
execution contract specifies not just a language but the exact library and API
|
||
function to use for that operation — chosen deterministically, with no LLM or
|
||
human in the loop for that decision.
|
||
|
||
---
|
||
|
||
## The Problem Being Solved
|
||
|
||
Current pipeline output (after Phase 5):
|
||
```
|
||
taskitem: "implement JSON serialization for UserRecord"
|
||
executionContract: { language: "C++", stepId: "serialize" }
|
||
```
|
||
|
||
Target pipeline output (after Phase 6):
|
||
```
|
||
taskitem: "implement JSON serialization for UserRecord"
|
||
executionContract: {
|
||
language: "C++",
|
||
library: "nlohmann_json",
|
||
operationDomain: "serialization.json",
|
||
preferredAPIs: ["nlohmann::json::dump()", "nlohmann::json::parse()"],
|
||
capabilityScore: 0.92,
|
||
avoidedLibraries: [{ name: "protobuf", reason: "no-pretty-print, slow-large-arrays" }],
|
||
justification: "nlohmann scores 0.92 vs protobuf 0.60 for serialization.json"
|
||
}
|
||
```
|
||
|
||
The SLM receives a pre-decided contract. Its only job: write syntax around a
|
||
pre-specified API call. Smallest possible decision space → highest possible
|
||
accuracy → closest to determinism.
|
||
|
||
---
|
||
|
||
## Architecture
|
||
|
||
### New Components
|
||
|
||
```
|
||
OperationTaxonomy — fixed registry of operation domains
|
||
anchored to what spec writers naturally write
|
||
e.g. serialization.json, compute.matrix, network.http.client,
|
||
io.stream.audio, crypto.hash, db.query.sql, ...
|
||
|
||
LibraryCapabilityRecord — per (library, operationDomain):
|
||
{ score: float, preferredAPIs: [], knownWeaknesses: [],
|
||
notApplicable: bool, annotatedAt: date, source: str }
|
||
|
||
LibraryCapabilityLedger — registry of LibraryCapabilityRecords
|
||
keyed by (libraryId, operationDomain)
|
||
one-time annotation, deterministic reads forever
|
||
|
||
PerTaskLibrarySelector — given (normalizedRequirements, availableLibraries):
|
||
for each operation domain in the task:
|
||
rank available libraries by score
|
||
return LibrarySelection { library, domain, score,
|
||
preferredAPIs, avoidedLibraries }
|
||
fully deterministic: sort by score, pick max, log justification
|
||
|
||
LibrarySymbolAdvisor — given (library, operationDomain):
|
||
return specific API symbols for that operation
|
||
seeded from: LSP symbol index, pre-built catalog, or both
|
||
deterministic lookup, no inference required
|
||
|
||
[Extended] generate_taskitems — execution contract gains:
|
||
selectedLibrary, operationDomain, capabilityScore,
|
||
preferredAPIs, avoidedLibraries, justification
|
||
```
|
||
|
||
### Selection Pipeline (fully deterministic after annotation)
|
||
|
||
```
|
||
spec text
|
||
→ architect_intake normalize requirements, extract operation domains
|
||
→ language fitness scorer pick language per operation (Phase 1, complete)
|
||
→ library capability ledger pick library per operation domain
|
||
(deterministic: max score lookup)
|
||
→ library symbol advisor pick specific API per operation
|
||
(deterministic: catalog lookup)
|
||
→ generate_taskitems execution contract: language + library + APIs
|
||
→ SLM / code generator fills in syntax only — no decisions required
|
||
```
|
||
|
||
---
|
||
|
||
## Operation Taxonomy Design Principles
|
||
|
||
1. **Anchored to spec language** — domains are what a spec writer writes, not what
|
||
a library author writes. "serialize to JSON" not "invoke jackson.ObjectMapper".
|
||
|
||
2. **Domain granularity** — coarser than individual functions, finer than "IO":
|
||
- `serialization.json` ≠ `serialization.binary` ≠ `serialization.xml`
|
||
- `compute.matrix` ≠ `compute.ml.training` ≠ `compute.ml.inference`
|
||
- `network.http.client` ≠ `network.http.server` ≠ `network.websocket`
|
||
|
||
3. **Hierarchical** — `serialization.json` is a child of `serialization`,
|
||
enabling fallback scoring when a library annotates at the parent level.
|
||
|
||
4. **Finite and versioned** — the taxonomy is a versioned registry, not free text.
|
||
New domains are added deliberately, not automatically.
|
||
|
||
---
|
||
|
||
## Library Capability Scoring Principles
|
||
|
||
- **Score range:** 0.00 (not applicable) to 1.00 (reference implementation)
|
||
- **Score meaning:**
|
||
- 0.00 — not applicable for this domain
|
||
- 0.60–0.75 — functional but with significant known weaknesses
|
||
- 0.76–0.90 — good, minor weaknesses or API ergonomic issues
|
||
- 0.91–1.00 — preferred reference for this domain
|
||
- **Annotation sources** (in priority order):
|
||
1. Formal benchmark data (throughput, latency, memory, correctness coverage)
|
||
2. Known CVE / correctness gaps from public vulnerability databases
|
||
3. API surface analysis (does a direct function exist for the operation?)
|
||
4. Community consensus from spec sheets / documentation
|
||
- **Annotation is one-time authoring** — the selector runs deterministically forever
|
||
- **Weaknesses are first-class** — a library can score 0.95 for binary serialization
|
||
and explicitly 0.60 for JSON serialization; both entries exist in the ledger
|
||
|
||
---
|
||
|
||
## Sprint Breakdown
|
||
|
||
### Sprint 286 — OperationTaxonomy + LibraryCapabilityLedger (steps 1958–1962)
|
||
|
||
| Step | Component | Description |
|
||
|------|-----------|-------------|
|
||
| 1958 | OperationDomain | Enum/registry of operation domains with hierarchy; lookup, parent traversal |
|
||
| 1959 | LibraryCapabilityRecord | Struct + validator; score range enforcement, weakness list, API list |
|
||
| 1960 | LibraryCapabilityLedger | Registry keyed by (libraryId, operationDomain); get/set/has/score |
|
||
| 1961 | OperationTaxonomy | Full domain tree with 40+ initial entries; classify(text)→domain |
|
||
| 1962 | Integration | Seed ledger with 8 real libraries × their operation domains; verify lookups |
|
||
|
||
**Initial library set for seeding:**
|
||
- `nlohmann_json` — serialization.json: 0.92
|
||
- `protobuf` — serialization.binary: 0.95, serialization.json: 0.60
|
||
- `serde_json` (Rust) — serialization.json: 0.98
|
||
- `cublas` — compute.matrix: 0.99 (GPU), compute.matrix: 0.00 (CPU fallback)
|
||
- `thrust` — compute.parallel: 0.97
|
||
- `numpy` — compute.matrix: 0.95, compute.statistics: 0.93
|
||
- `tokio` (Rust) — network.async: 0.97, io.stream: 0.95
|
||
- `boost.asio` — network.async: 0.91, network.http.client: 0.78
|
||
|
||
---
|
||
|
||
### Sprint 287 — PerTaskLibrarySelector (steps 1963–1967)
|
||
|
||
| Step | Component | Description |
|
||
|------|-----------|-------------|
|
||
| 1963 | LibrarySelection | Struct: library, domain, score, preferredAPIs, avoidedLibraries, justification |
|
||
| 1964 | OperationClassifier | Classifies normalized requirement text → OperationDomain (deterministic) |
|
||
| 1965 | PerTaskLibrarySelector | Core selector: rank by score, return best + avoided list + justification |
|
||
| 1966 | SelectionJustificationLog | Structured log of every selection decision (auditable, deterministic replay) |
|
||
| 1967 | Integration | 8-task spec with mixed operations; verify correct library chosen per task |
|
||
|
||
---
|
||
|
||
### Sprint 288 — LibrarySymbolAdvisor (steps 1968–1972)
|
||
|
||
| Step | Component | Description |
|
||
|------|-----------|-------------|
|
||
| 1968 | LibrarySymbolRecord | Struct: functionName, signature, operationDomain, usagePattern, caveats |
|
||
| 1969 | LibrarySymbolCatalog | Registry of LibrarySymbolRecords; lookup by (library, operationDomain) |
|
||
| 1970 | LSPSymbolExtractor | Parses LSP completion/hover responses → LibrarySymbolRecords |
|
||
| 1971 | LibrarySymbolAdvisor | Given (library, operationDomain) → ranked symbol recommendations |
|
||
| 1972 | Integration | CUDA: advisee returns cublasSgemm for compute.matrix; nlohmann: json::parse |
|
||
|
||
---
|
||
|
||
### Sprint 289 — Integration into generate_taskitems (steps 1973–1977)
|
||
|
||
| Step | Component | Description |
|
||
|------|-----------|-------------|
|
||
| 1973 | EnrichedExecutionContract | Extends existing contract with library dispatch fields |
|
||
| 1974 | TaskitemLibraryAnnotator | Runs selector + advisor per taskitem; injects into contract |
|
||
| 1975 | AnnotatedTaskitemValidator | Validates library fields present when domain is known |
|
||
| 1976 | DispatchJustificationReport | Per-project report: every task → library choice → score → why |
|
||
| 1977 | Integration | Full pipeline: spec → intake → selector → advisor → enriched taskitems |
|
||
|
||
---
|
||
|
||
### Sprint 290 — CUDA End-to-End Proof (steps 1978–1982)
|
||
|
||
| Step | Component | Description |
|
||
|------|-----------|-------------|
|
||
| 1978 | CUDALibraryProfile | Full CUDA ledger entries: cublas, cufft, thrust, cudnn domains + scores |
|
||
| 1979 | CUDASymbolCatalog | 20+ CUDA API symbols with signatures and usage patterns |
|
||
| 1980 | ProjectStackDeclarator | Declares available libraries for a project (input to selector) |
|
||
| 1981 | ProofPipelineRunner | Runs full Phase 6 pipeline on a CUDA project spec |
|
||
| 1982 | Integration | Spec: "GPU matrix multiply + JSON result serialization"
|
||
Output: cublasSgemm for compute.matrix, nlohmann for serialization.json
|
||
Sprint290IntegrationSummary |
|
||
|
||
---
|
||
|
||
## What Changes for the SLM After Phase 6
|
||
|
||
**Before Phase 6** — SLM decides:
|
||
- What library to use (reasoning required)
|
||
- What API to call (knowledge required)
|
||
- Which edge cases the library handles (judgment required)
|
||
|
||
**After Phase 6** — SLM receives:
|
||
```
|
||
language: C++
|
||
library: cublas
|
||
operationDomain: compute.matrix
|
||
preferredAPIs: ["cublasSgemm(handle, CUBLAS_OP_N, CUBLAS_OP_N, m, n, k, &alpha, A, lda, B, ldb, &beta, C, ldc)"]
|
||
knownWeaknesses: ["requires-device-memory-management", "no-automatic-handle-lifecycle"]
|
||
```
|
||
|
||
SLM job: write the function body that calls this API with these parameters.
|
||
No reasoning. No library knowledge. Syntax only.
|
||
|
||
---
|
||
|
||
## Impact on Training Data
|
||
|
||
After Phase 6 is live:
|
||
- **Retire:** `mcp_calls_good.jsonl`, `mcp_call_quality.jsonl` — these teach
|
||
the SLM to make decisions that are now deterministic. Training on them teaches
|
||
wrong behavior.
|
||
- **Archive:** `mcp_calls_bad.jsonl` — failure modes are architecture-independent.
|
||
- **Keep:** `taskitem_pipeline_runs.jsonl` — pipeline structure signal remains valid.
|
||
- **Keep:** `datasets/` benchmark specs — reusable eval sets regardless of architecture.
|
||
|
||
New training data target post-Phase 6: syntax-fill examples where the contract
|
||
is fully specified and the SLM is evaluated only on correctness of the generated
|
||
function body — not on the library/API decision.
|
||
|
||
---
|
||
|
||
## Success Criteria for Phase 6
|
||
|
||
1. Given a project spec with a declared library stack, every taskitem execution
|
||
contract contains: `selectedLibrary`, `operationDomain`, `capabilityScore`,
|
||
`preferredAPIs`, `avoidedLibraries`, `justification`
|
||
2. Selection is deterministic: same spec + same ledger = same output, always
|
||
3. No LLM or human consulted during selection
|
||
4. CUDA project proof: "GPU matrix multiply" → `cublasSgemm`, not a kernel loop
|
||
5. Mixed-stack proof: project uses protobuf for IPC but `serialization.json` task
|
||
correctly selects `nlohmann_json` over protobuf (0.92 vs 0.60)
|