Files
whetstone_DSL/LO_RA_HANDOFF_2026-02-25.md

113 lines
5.2 KiB
Markdown
Raw Permalink Normal View History

# LoRA Training Handoff - 2026-02-25
## Scope
This note captures what failed on 2026-02-25 (UTC), what was fixed, what is still failing, and how to continue without re-discovering context.
## User intent
- Train LoRA for Qwen 7B instruct (`Qwen/Qwen2.5-Coder-7B-Instruct`) after 14B OOM issues.
- Use persistent process (tmux/nohup) so terminal close does not kill training.
- Report-only monitoring (no watchdog restarts).
- Run short 20-step benchmark chunks first, then scale.
## Major failure timeline (2026-02-25)
1. 14B path OOM / memory-fit instability
- Symptom: repeated OOM and/or load failures with 14B under available VRAM.
- Note: user system RAM availability does not automatically solve VRAM-only allocations; many failures were from GPU-side allocation limits and bnb dispatch constraints.
2. Sandbox vs host GPU visibility mismatch
- Symptom: `nvidia-smi` worked in user terminal but failed intermittently in sandbox runs.
- Impact: false negatives while checking GPU status.
- Action: moved critical launch/monitor commands to host-level escalated execution.
3. Process persistence confusion
- Symptom: runs appeared to stop when terminals closed.
- Root cause: non-persistent shell sessions for some launches.
- Action: switched to tmux/nohup-style launchers.
4. Quantization path failures
- Int8 failure seen:
- `TypeError: Int8Params.__new__() got an unexpected keyword argument '_is_hf_initialized'`
- 4-bit CPU/disk dispatch failure seen:
- `ValueError: Some modules are dispatched on the CPU or the disk...` (from `quantizer_bnb_4bit.validate_environment`)
- Additional OOM during load seen in some offload variants:
- `torch.OutOfMemoryError... Tried to allocate 26.00 MiB...`
5. Transformers/TRL API drift issues
- Fixed earlier:
- `TrainingArguments.__init__() got an unexpected keyword argument 'evaluation_strategy'`
- Script updated to detect `eval_strategy` vs `evaluation_strategy`.
- New failure discovered later (after model load + dataset map):
- `TypeError: SFTTrainer.__init__() got an unexpected keyword argument 'max_seq_length'`
- This was the immediate reason a long load appeared to "stall then die".
## Concrete fix set applied today
### 1) TRL SFTTrainer compatibility fix
File: `tools/mcp/train_quality_lora.py`
- Switched trainer arg construction to TRL `SFTConfig`.
- Added runtime signature checks for `eval_strategy`/`evaluation_strategy`.
- Moved max sequence handling to `SFTConfig(max_length=...)`.
- Added conditional injection of `max_seq_length` only if current `SFTTrainer` accepts it.
- Result: avoids crash at `SFTTrainer(... max_seq_length=...)` on this installed TRL version.
### 2) Checkpoint cadence for resumability in short chunks
File: `tools/mcp/run_qwen14b_quality_chunk.sh`
- Added auto override for `SAVE_STEPS_OVERRIDE` based on chunk size:
- `save_steps = max(1, CHUNK_STEPS/4)` if not explicitly set.
- Passes `SAVE_STEPS_OVERRIDE` down into training launcher.
- Result: 20-step benchmark chunks now checkpoint during run (instead of default save every 200, which gave nothing resumable for short runs).
## Current run state at handoff
- A fresh forensic run was started in tmux session: `lora7b_forensic`.
- Latest forensic run id observed: `20260225T013638Z`.
- Files:
- `/tmp/qwen7b_forensic_20260225T013638Z.log`
- `/tmp/qwen7b_forensic_20260225T013638Z.exit`
- `/tmp/qwen7b_forensic_20260225T013638Z.heartbeat.log`
- At last check:
- trainer process existed (`train_quality_lora.py ... --save-steps 5 --max-steps 20 ...`).
- run was in weight loading progress (not yet confirmed fully complete in this final check).
## Key operational truths to preserve
- After crash/exit, model will NOT remain loaded. Resume is checkpoint-based, not in-memory continuation.
- `--auto-resume` only helps if checkpoint directories exist in output dir.
- For short benchmark runs, save cadence must be low enough (now addressed).
- Use host-level GPU checks if sandbox gives contradictory `nvidia-smi` behavior.
## Commands for next agent (copy/paste)
1. Verify session/process:
```bash
tmux has-session -t lora7b_forensic && echo alive || echo dead
pgrep -af "train_quality_lora.py --base-model Qwen/Qwen2.5-Coder-7B-Instruct|run_qwen7b_forensic_once.sh"
```
2. Inspect latest forensic artifacts:
```bash
ls -lt /tmp/qwen7b_forensic_* | head -n 20
tail -n 120 /tmp/qwen7b_forensic_20260225T013638Z.log
cat /tmp/qwen7b_forensic_20260225T013638Z.exit
tail -n 40 /tmp/qwen7b_forensic_20260225T013638Z.heartbeat.log
```
3. Check whether resumable checkpoints now exist:
```bash
ls -lt /home/bill/Documents/CLionProjects/whetstone_DSL/training_data/lora/adapters/qwen2.5-coder-7b-quality-lora
ls -d /home/bill/Documents/CLionProjects/whetstone_DSL/training_data/lora/adapters/qwen2.5-coder-7b-quality-lora/checkpoint-* 2>/dev/null | tail
```
4. If re-run is needed (same forensic path):
```bash
cd /home/bill/Documents/CLionProjects/whetstone_DSL
./tools/mcp/run_qwen7b_forensic_once.sh
```
## Known open risks
- bitsandbytes/accelerate/transformers/trl version mismatches may still surface at later phases.
- 4-bit + strict offload combinations can still trigger device-map validation errors depending on runtime environment.
- Even with checkpointing fixed, wall-clock can still be dominated by repeated model loads if runs keep failing before stable stepping.