7.4 KiB
UncheckedIO — Development Roadmap
Goal: Production-grade, multi-GPU zero-copy ingestion pipeline for LLM providers. Current state: Working proof-of-concept on CUDA (gpu-zero branch). PostgreSQL→Arrow open source on main.
Phase 1 — Fix Correctness (gpu-zero branch)
The GPU feature works in benchmark conditions but has four bugs that would surface immediately in production. Fix these before anything else.
1.1 Dynamic buffer allocation
Problem: max_elements = 4096 * 512 is hardcoded at init time. Modern LLMs have 32K–128K context windows. A single request at 32K tokens + any real batch size hits this limit.
Fix: Allocate gpu_buffer and host_buffer per-call based on actual batch_size * max_seq_len, or use a growable strategy that reallocates when the current batch exceeds the existing buffer size.
1.2 Buffer isolation (silent data corruption bug)
Problem: Every call to encode_batch overwrites the same CudaSlice. The DLPack capsule hands PyTorch a pointer to that same memory. A caller holding a tensor from call N gets its data silently overwritten by call N+1.
Fix: Double-buffer (ping-pong between two allocations) or allocate a fresh device buffer per call and transfer ownership to the DLPack capsule's deleter so CUDA frees it when PyTorch releases the tensor.
1.3 Thread safety
Problem: encode_batch takes &mut self. The TokenizerEngine Python class cannot be called from multiple threads simultaneously. Every production serving engine (vLLM, TGI) is multithreaded.
Fix: Wrap PinnedBatcher in an Arc<Mutex<>> or, better, create a EnginePool that holds N instances and checks one out per call. Pool size = CPU core count is a reasonable default.
1.4 Variable sequence length + padding
Problem: seq_len = encodings[0].len() assumes all sequences in the batch have identical length. The tokenizers crate doesn't apply padding by default, so real variable-length inputs produce an incorrect 2D tensor shape.
Fix: Apply padding to max(encodings[i].len()) within the batch, fill with the tokenizer's pad_id, and use the actual padded length as seq_len.
Phase 2 — Production Features
2.1 vLLM integration
vLLM is the dominant open-source LLM serving engine. Make unchecked_io a drop-in tokenization backend for it.
- Implement the
TokenizerGroupprotocol vLLM expects - The continuous batching scheduler sends variable-length requests asynchronously — this is where Phase 1 fixes must be solid
- Target: a 3-line config change in vLLM to switch to the UncheckedIO backend
- Deliverable:
examples/vllm_integration.pyshowing the swap
2.2 Streaming / chunked output
Current: The entire batch is tokenized and DMA'd in one blocking call. Target: Yield Arrow RecordBatches or DLPack tensors incrementally as tokens are ready, so the model can begin prefill before the full batch is staged.
2.3 Attention mask export
LLM inference requires both input_ids and attention_mask. Currently only token IDs are exported. Add a second DLPack capsule for the mask, or pack both into a named dict capsule.
2.4 Configurable max context
Replace the hardcoded constants with constructor arguments:
engine = TokenizerEngine(
tokenizer_path="tokenizer.json",
max_seq_len=32768,
max_batch_size=64,
pool_size=8
)
Phase 3 — Multi-GPU Support
3.1 ROCm / HIP (AMD) — highest priority
AMD's MI300X has 192GB memory vs the H100's 80GB and providers are actively evaluating it to escape NVIDIA pricing. HIP is the path of least resistance — function names are nearly 1:1 with CUDA (hipMemHostRegister, hipMemcpyHtoD, hipMemHostUnregister).
- Add
rocmfeature flag inCargo.toml - Create
src/backends/rocm.rsimplementing the sameGpuBackendtrait - Use AMD's
hip-sysorrocm-sysRust crate - The DLPack export layer stays identical — device_type
10for ROCm vs2for CUDA
Refactor needed first: Extract a GpuBackend trait from the current CUDA-specific code:
trait GpuBackend {
fn allocate_pinned(capacity: usize) -> Result<PinnedBuffer>;
fn copy_to_device(src: &[i64], dst: &mut DeviceBuffer) -> Result<()>;
fn device_ptr(buf: &DeviceBuffer) -> u64;
fn dlpack_device_type() -> i32;
}
3.2 OpenCL (broad coverage)
Covers AMD, Intel, older NVIDIA, Apple (via Metal-OpenCL bridge). Lower peak performance than native CUDA/HIP but maximum hardware coverage. Use the opencl3 Rust crate.
- Add
openclfeature flag - Useful for providers running mixed-GPU clusters
3.3 Intel Gaudi / Level Zero
Intel's Gaudi 2/3 accelerators are gaining traction for training and are available on AWS and Azure. Intel Level Zero is the low-level API equivalent to CUDA driver API. Lower priority than ROCm but worth tracking.
Phase 4 — Benchmark Suite
The current benchmark_gpu.py is good for what it tests but synthetic. A buyer will run their own benchmarks — make it easy and cover real scenarios.
4.1 Scenarios to cover
- Short sequences: 128, 256 tokens (BERT-style)
- Medium sequences: 2048, 4096 tokens (typical chat)
- Long sequences: 16K, 32K tokens (document/RAG workloads)
- Variable length batches (realistic distribution, not identical strings)
- Concurrent throughput: simulate N async callers
4.2 Tokenizers to cover
- BERT (
bert-base-uncased) — already done - LLaMA 3 tokenizer (BPE, more representative of modern LLMs)
- GPT-2 tokenizer (common baseline)
4.3 Hardware targets
- Local RTX 3060 (done, 4.7x)
- Cloud H100 SXM (~$2 on Lambda Labs, 30 min run) — do this first, it changes the pitch
- AMD MI300X (after Phase 3.1)
4.4 What to measure
- Tokens/second throughput
- Latency per batch (p50, p95, p99)
- GPU utilization % (the thing providers care about)
- Memory bandwidth saturation
Phase 5 — Business Infrastructure
5.1 Wyoming LLC formation
- Single-member LLC, Wyoming (strongest privacy, no public member disclosure)
- LLC owns all IP
- Registered agent handles public address
- Done before any contract is signed
5.2 License structure
- Open core: PostgreSQL→Arrow stays Apache 2.0
- GPU features: commercial license
- Pricing model: annual per-cluster license, priced against GPU utilization savings
- License enforcement: honor-system enterprise contracts initially, not technical DRM
5.3 Sales collateral
- 1-page technical brief (problem / mechanism / numbers / what's included)
- Self-contained benchmark the prospect can run on their own hardware
- Reference architecture diagram showing where this sits in a vLLM serving stack
Branch Strategy
| Branch | Purpose |
|---|---|
main |
Open source PostgreSQL→Arrow, Apache 2.0, public |
gpu-zero |
GPU feature development, private |
rocm |
AMD port (Phase 3.1) |
release/x.y |
Tagged releases for licensing |
The gpu feature is compiled in only with --features gpu. The open source build never includes GPU code.
Immediate Next Steps (in order)
- Fix 1.1 (dynamic buffer) — unblocks everything else
- Fix 1.2 (buffer isolation) — correctness before benchmarking
- Fix 1.3 (thread safety) — required for any real serving test
- Fix 1.4 (padding) — correctness
- Spend $2, run benchmark on cloud H100 — this one number changes the pitch
- Phase 2.1 vLLM integration — this is the demo that sells the concept
- Phase 3.1 ROCm — this is the differentiator