Files
whetstone_DSL/LO_RA_HANDOFF_2026-02-25.md

5.2 KiB

LoRA Training Handoff - 2026-02-25

Scope

This note captures what failed on 2026-02-25 (UTC), what was fixed, what is still failing, and how to continue without re-discovering context.

User intent

  • Train LoRA for Qwen 7B instruct (Qwen/Qwen2.5-Coder-7B-Instruct) after 14B OOM issues.
  • Use persistent process (tmux/nohup) so terminal close does not kill training.
  • Report-only monitoring (no watchdog restarts).
  • Run short 20-step benchmark chunks first, then scale.

Major failure timeline (2026-02-25)

  1. 14B path OOM / memory-fit instability
  • Symptom: repeated OOM and/or load failures with 14B under available VRAM.
  • Note: user system RAM availability does not automatically solve VRAM-only allocations; many failures were from GPU-side allocation limits and bnb dispatch constraints.
  1. Sandbox vs host GPU visibility mismatch
  • Symptom: nvidia-smi worked in user terminal but failed intermittently in sandbox runs.
  • Impact: false negatives while checking GPU status.
  • Action: moved critical launch/monitor commands to host-level escalated execution.
  1. Process persistence confusion
  • Symptom: runs appeared to stop when terminals closed.
  • Root cause: non-persistent shell sessions for some launches.
  • Action: switched to tmux/nohup-style launchers.
  1. Quantization path failures
  • Int8 failure seen:
    • TypeError: Int8Params.__new__() got an unexpected keyword argument '_is_hf_initialized'
  • 4-bit CPU/disk dispatch failure seen:
    • ValueError: Some modules are dispatched on the CPU or the disk... (from quantizer_bnb_4bit.validate_environment)
  • Additional OOM during load seen in some offload variants:
    • torch.OutOfMemoryError... Tried to allocate 26.00 MiB...
  1. Transformers/TRL API drift issues
  • Fixed earlier:
    • TrainingArguments.__init__() got an unexpected keyword argument 'evaluation_strategy'
    • Script updated to detect eval_strategy vs evaluation_strategy.
  • New failure discovered later (after model load + dataset map):
    • TypeError: SFTTrainer.__init__() got an unexpected keyword argument 'max_seq_length'
    • This was the immediate reason a long load appeared to "stall then die".

Concrete fix set applied today

1) TRL SFTTrainer compatibility fix

File: tools/mcp/train_quality_lora.py

  • Switched trainer arg construction to TRL SFTConfig.
  • Added runtime signature checks for eval_strategy/evaluation_strategy.
  • Moved max sequence handling to SFTConfig(max_length=...).
  • Added conditional injection of max_seq_length only if current SFTTrainer accepts it.
  • Result: avoids crash at SFTTrainer(... max_seq_length=...) on this installed TRL version.

2) Checkpoint cadence for resumability in short chunks

File: tools/mcp/run_qwen14b_quality_chunk.sh

  • Added auto override for SAVE_STEPS_OVERRIDE based on chunk size:
    • save_steps = max(1, CHUNK_STEPS/4) if not explicitly set.
  • Passes SAVE_STEPS_OVERRIDE down into training launcher.
  • Result: 20-step benchmark chunks now checkpoint during run (instead of default save every 200, which gave nothing resumable for short runs).

Current run state at handoff

  • A fresh forensic run was started in tmux session: lora7b_forensic.
  • Latest forensic run id observed: 20260225T013638Z.
  • Files:
    • /tmp/qwen7b_forensic_20260225T013638Z.log
    • /tmp/qwen7b_forensic_20260225T013638Z.exit
    • /tmp/qwen7b_forensic_20260225T013638Z.heartbeat.log
  • At last check:
    • trainer process existed (train_quality_lora.py ... --save-steps 5 --max-steps 20 ...).
    • run was in weight loading progress (not yet confirmed fully complete in this final check).

Key operational truths to preserve

  • After crash/exit, model will NOT remain loaded. Resume is checkpoint-based, not in-memory continuation.
  • --auto-resume only helps if checkpoint directories exist in output dir.
  • For short benchmark runs, save cadence must be low enough (now addressed).
  • Use host-level GPU checks if sandbox gives contradictory nvidia-smi behavior.

Commands for next agent (copy/paste)

  1. Verify session/process:
tmux has-session -t lora7b_forensic && echo alive || echo dead
pgrep -af "train_quality_lora.py --base-model Qwen/Qwen2.5-Coder-7B-Instruct|run_qwen7b_forensic_once.sh"
  1. Inspect latest forensic artifacts:
ls -lt /tmp/qwen7b_forensic_* | head -n 20
tail -n 120 /tmp/qwen7b_forensic_20260225T013638Z.log
cat /tmp/qwen7b_forensic_20260225T013638Z.exit
tail -n 40 /tmp/qwen7b_forensic_20260225T013638Z.heartbeat.log
  1. Check whether resumable checkpoints now exist:
ls -lt /home/bill/Documents/CLionProjects/whetstone_DSL/training_data/lora/adapters/qwen2.5-coder-7b-quality-lora
ls -d /home/bill/Documents/CLionProjects/whetstone_DSL/training_data/lora/adapters/qwen2.5-coder-7b-quality-lora/checkpoint-* 2>/dev/null | tail
  1. If re-run is needed (same forensic path):
cd /home/bill/Documents/CLionProjects/whetstone_DSL
./tools/mcp/run_qwen7b_forensic_once.sh

Known open risks

  • bitsandbytes/accelerate/transformers/trl version mismatches may still surface at later phases.
  • 4-bit + strict offload combinations can still trigger device-map validation errors depending on runtime environment.
  • Even with checkpointing fixed, wall-clock can still be dominated by repeated model loads if runs keep failing before stable stepping.