# Handoff — 2026-04-21 ## What Was Done This Session Fixed all four Phase 1 correctness bugs. Code compiles clean on temp-ubuntu with CUDA 13.1. --- ## Phase 1 Bugs — ALL FIXED ### 1.1 Dynamic buffer allocation — DONE **Old behavior:** `max_elements = 4096 * 512` hardcoded in `PinnedBatcher::new`. Any batch exceeding ~2M tokens returned an error. **Fix:** Removed `gpu_buffer`, `host_buffer`, and `max_elements` from the `PinnedBatcher` struct entirely. Both buffers are now allocated per-call based on actual `batch_size * seq_len`. ### 1.2 Buffer isolation (DLPack overwrite) — DONE **Old behavior:** Every call overwrote the same `CudaSlice`. Any tensor from call N got corrupted by call N+1. **Fix:** Each call allocates its own device buffer and wraps it in a `DLPackContext` struct stored in `DLManagedTensor.manager_ctx`. The deleter frees both the device memory (by dropping `CudaSlice`) and the shape array when PyTorch releases the tensor. Each returned tensor now owns its allocation. ### 1.3 Thread safety — DONE **Old behavior:** `batch_encode_to_gpu` took `&mut self`. `TokenizerEngine.encode_batch` took `&mut self`. Not usable from multiple Python threads. **Fix:** Since `PinnedBatcher` no longer has mutable shared state (buffers are per-call), `batch_encode_to_gpu` now takes `&self`. `TokenizerEngine` wraps `Arc` instead of owning it directly. `encode_batch` takes `&self`. Concurrent calls from Python threads are safe. ### 1.4 Variable sequence length + padding — DONE **Old behavior:** `seq_len = encodings[0].len()` — assumed uniform length. Variable-length batches produced wrong tensor shape. **Fix:** `seq_len = encodings.iter().map(|e| e.len()).max()`. Host buffer prefilled with `pad_id` (from `tokenizer.get_padding()`, fallback to 0). Rayon fill loop only writes actual tokens; shorter sequences retain the pad value. --- ## Build Issue Discovered and Resolved `cudarc 0.11.9` panics at build time on CUDA 13.1 (only knows 12.5 and below). This machine has CUDA 13.1. **Fix:** Changed the `gpu` feature in `Cargo.toml` from `cudarc/cuda-version-from-build-system` to `cudarc/cuda-12050`. This hard-pins to CUDA 12.5 headers and skips the `nvcc` version check. CUDA 13.x is backward-compatible with 12.5 driver API at runtime. If you ever upgrade cudarc to 0.12+, revert this change and re-test. --- ## Files Changed | File | What Changed | |------|-------------| | `src/llm.rs` | Full rewrite of `PinnedBatcher` — removed static buffers, per-call allocation, `DLPackContext` for ownership, `&self` throughout | | `src/lib.rs` | `TokenizerEngine.inner` changed to `Arc`, `encode_batch` takes `&self`, added `use std::sync::Arc` | | `Cargo.toml` | `gpu` feature: `cuda-version-from-build-system` → `cuda-12050` | --- ## This Week — What's Left (in order) ### Today / Tomorrow (Tuesday 2026-04-22) - [ ] Rent H100 on Lambda Labs (~$2 for 30 min) - [ ] Run `benchmark_gpu.py` under `nsys profile --stats=true python benchmark_gpu.py` - [ ] Capture Nsight `.nsys-rep` file and download it - [ ] Note the speedup number — this is the headline for Thursday's talk **Still need:** H100 setup script (a single paste-into-terminal script for the cloud instance). Needs: Rust install, maturin build with `--features gpu`, nsys profile run, output capture. This was not written last session. ### Wednesday (2026-04-23) - [ ] Update `docs/talk-10min.md` Slide 6 with the H100 result - [ ] Open Nsight Systems GUI locally, load the `.nsys-rep` file, screenshot the timeline showing both pipelines ### Thursday (2026-04-24) - [ ] Talk to IT/security group (10 min) - [ ] If someone gives you a contact: do not show GPU source code without NDA signed first; first call is discovery --- ## Next Session Priorities (after Thursday) 1. **Write H100 setup script** — single paste-into-terminal script (see above) 2. **vLLM integration** — Phase 2.1, this is the demo that sells the concept 3. **NM anonymous LLC** — Wyoming, single-member, before any contract is signed 4. **C++ version scoping** — after Phase 1 solid, assess Whetstone transpilation of `llm.rs` --- ## Key Technical Context - Repo: `/home/bill/Documents/unchecked-io/`, on `gpu-zero` branch - Build: `maturin develop --release --features gpu` (maturin not installed on temp-ubuntu — install with `pip install maturin` first) - `cargo check --features gpu` confirms clean compile - The four Phase 1 bugs are fixed and verified to compile - `src/llm.rs` is the GPU core. `src/lib.rs` is the Python-facing wrapper. - `chrono-tz` pinned at `=0.10.4` — do not change - `gpu-zero` branch is private; `main` is open source Apache 2.0