3.0 KiB
3.0 KiB
Multi-Model LoRA Data Playbook
Goal
Build a reusable dataset where failures come from multiple models and good examples come from validated pipeline outputs, then rebalance into a training-ready JSONL.
Required metadata per record
Every record should carry:
model(for exampleqwen2.5-coder:14b,qwen2.5-coder:7b)label(goodorbad)interface_mode(json_loopornative_tool_calling)source(for examplepipeline_run,stage_labeled,mcp_tool_response)
Collection flow
- Run collection loops per model.
- Export malformed and quality datasets after each run.
- Keep run manifests for token/cycle diagnostics.
- Build one balanced multimodel output before training.
Suggested run strategy
- Keep one primary target model (for example 14B) as highest share of data.
- Add 1-2 secondary models for failure diversity.
- Avoid overfitting to one repeated failure signature.
Commands
0) Good-data hybrid flow (deterministic + optional stronger-LLM review)
# Step A: deterministic good-data growth from run specs (same corpora as bad-data loop by default)
./tools/mcp/generate_good_from_run_specs.sh
# Step B: split good records into high-confidence and review queue
./tools/mcp/build_good_hybrid_dataset.sh
# One-command orchestrator (A+B)
./tools/mcp/run_good_data_hybrid.sh
# After stronger-LLM review creates approved-good JSONL:
REVIEW_APPROVED_FILE=training_data/lora/mcp_calls_good_review_approved.jsonl \
./tools/mcp/run_good_data_hybrid.sh
1) Per-run exports
./tools/mcp/export_malformed_mcp_calls_for_lora.sh logs/taskitem_runs/<bridge_run_dir>
./tools/mcp/export_mcp_call_quality_for_lora.sh logs/taskitem_runs/<good_pipeline_run_dir> logs/taskitem_runs/<bridge_run_dir>
2) Build balanced multimodel dataset
BAD_TO_GOOD_RATIO=1.0 \
MAX_PER_MODEL_PER_LABEL=250000 \
REQUIRE_MODEL_METADATA=0 \
DEFAULT_MODEL_TAG=unknown_model \
DEFAULT_INTERFACE_MODE=json_loop \
./tools/mcp/build_multimodel_lora_dataset.sh
Outputs:
training_data/lora/mcp_multimodel_balanced.jsonltraining_data/lora/mcp_multimodel_balanced.summary.json
3) Create leakage-safe train/val/test split (grouped by run/spec family)
./tools/mcp/split_lora_dataset_by_group.sh
Outputs:
training_data/lora/splits/mcp_lora.train.jsonltraining_data/lora/splits/mcp_lora.val.jsonltraining_data/lora/splits/mcp_lora.test.jsonltraining_data/lora/splits/mcp_lora.split_summary.json
Quality gates before training
- Confirm
selected_goodandselected_badcounts are both non-trivial. - Confirm multiple models appear in
by_model_selected. - Confirm source diversity in
by_source_selected. - Spot-check for infra-only failures; keep operational noise low.
For future agents
When collecting new runs:
- Do not overwrite raw files.
- Rebuild balanced output from raw files each time.
- Record all run IDs and model tags in commit notes.
When training:
- Train separate LoRAs per base model family/size.
- Reuse the same balanced dataset recipe; do not reuse adapters across base sizes.