Files
unchecked-io/HANDOFF-2026-04-21.md
Bill Holcombe dc8494bb2a Fix all four Phase 1 correctness bugs (dynamic buffer, DLPack ownership, thread safety, padding)
- 1.1: Remove hardcoded 4096*512 buffer from PinnedBatcher; allocate host+device per-call
- 1.2: Wrap CudaSlice in DLPackContext owned by DLManagedTensor.manager_ctx; deleter frees
  device memory when PyTorch releases tensor — no more silent overwrite across calls
- 1.3: batch_encode_to_gpu now takes &self (no mutable shared state); TokenizerEngine wraps
  Arc<PinnedBatcher> and encode_batch takes &self — safe for concurrent Python threads
- 1.4: seq_len = max across all encodings; host buffer prefilled with pad_id before token write

Also: switch cudarc gpu feature from cuda-version-from-build-system to cuda-12050 to fix
build on machines with CUDA 13.x (cudarc 0.11.9 only knows up to 12.5)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-21 14:10:31 -06:00

4.6 KiB

Handoff — 2026-04-21

What Was Done This Session

Fixed all four Phase 1 correctness bugs. Code compiles clean on temp-ubuntu with CUDA 13.1.


Phase 1 Bugs — ALL FIXED

1.1 Dynamic buffer allocation — DONE

Old behavior: max_elements = 4096 * 512 hardcoded in PinnedBatcher::new. Any batch exceeding ~2M tokens returned an error. Fix: Removed gpu_buffer, host_buffer, and max_elements from the PinnedBatcher struct entirely. Both buffers are now allocated per-call based on actual batch_size * seq_len.

1.2 Buffer isolation (DLPack overwrite) — DONE

Old behavior: Every call overwrote the same CudaSlice. Any tensor from call N got corrupted by call N+1. Fix: Each call allocates its own device buffer and wraps it in a DLPackContext struct stored in DLManagedTensor.manager_ctx. The deleter frees both the device memory (by dropping CudaSlice<i64>) and the shape array when PyTorch releases the tensor. Each returned tensor now owns its allocation.

1.3 Thread safety — DONE

Old behavior: batch_encode_to_gpu took &mut self. TokenizerEngine.encode_batch took &mut self. Not usable from multiple Python threads. Fix: Since PinnedBatcher no longer has mutable shared state (buffers are per-call), batch_encode_to_gpu now takes &self. TokenizerEngine wraps Arc<PinnedBatcher> instead of owning it directly. encode_batch takes &self. Concurrent calls from Python threads are safe.

1.4 Variable sequence length + padding — DONE

Old behavior: seq_len = encodings[0].len() — assumed uniform length. Variable-length batches produced wrong tensor shape. Fix: seq_len = encodings.iter().map(|e| e.len()).max(). Host buffer prefilled with pad_id (from tokenizer.get_padding(), fallback to 0). Rayon fill loop only writes actual tokens; shorter sequences retain the pad value.


Build Issue Discovered and Resolved

cudarc 0.11.9 panics at build time on CUDA 13.1 (only knows 12.5 and below). This machine has CUDA 13.1.

Fix: Changed the gpu feature in Cargo.toml from cudarc/cuda-version-from-build-system to cudarc/cuda-12050. This hard-pins to CUDA 12.5 headers and skips the nvcc version check. CUDA 13.x is backward-compatible with 12.5 driver API at runtime. If you ever upgrade cudarc to 0.12+, revert this change and re-test.


Files Changed

File What Changed
src/llm.rs Full rewrite of PinnedBatcher — removed static buffers, per-call allocation, DLPackContext for ownership, &self throughout
src/lib.rs TokenizerEngine.inner changed to Arc<PinnedBatcher>, encode_batch takes &self, added use std::sync::Arc
Cargo.toml gpu feature: cuda-version-from-build-systemcuda-12050

This Week — What's Left (in order)

Today / Tomorrow (Tuesday 2026-04-22)

  • Rent H100 on Lambda Labs (~$2 for 30 min)
  • Run benchmark_gpu.py under nsys profile --stats=true python benchmark_gpu.py
  • Capture Nsight .nsys-rep file and download it
  • Note the speedup number — this is the headline for Thursday's talk

Still need: H100 setup script (a single paste-into-terminal script for the cloud instance). Needs: Rust install, maturin build with --features gpu, nsys profile run, output capture. This was not written last session.

Wednesday (2026-04-23)

  • Update docs/talk-10min.md Slide 6 with the H100 result
  • Open Nsight Systems GUI locally, load the .nsys-rep file, screenshot the timeline showing both pipelines

Thursday (2026-04-24)

  • Talk to IT/security group (10 min)
  • If someone gives you a contact: do not show GPU source code without NDA signed first; first call is discovery

Next Session Priorities (after Thursday)

  1. Write H100 setup script — single paste-into-terminal script (see above)
  2. vLLM integration — Phase 2.1, this is the demo that sells the concept
  3. NM anonymous LLC — Wyoming, single-member, before any contract is signed
  4. C++ version scoping — after Phase 1 solid, assess Whetstone transpilation of llm.rs

Key Technical Context

  • Repo: /home/bill/Documents/unchecked-io/, on gpu-zero branch
  • Build: maturin develop --release --features gpu (maturin not installed on temp-ubuntu — install with pip install maturin first)
  • cargo check --features gpu confirms clean compile
  • The four Phase 1 bugs are fixed and verified to compile
  • src/llm.rs is the GPU core. src/lib.rs is the Python-facing wrapper.
  • chrono-tz pinned at =0.10.4 — do not change
  • gpu-zero branch is private; main is open source Apache 2.0