3.8 KiB
3.8 KiB
Agent Notes 2026-02-23
Use this file as the primary handoff for restarting SLM data runs tomorrow.
Session summary
- Expanded data generation away from sprint plans to run specs and project fixtures.
- Ran long novelty churn cycles with heartbeat telemetry and category-balanced rebuilds.
- Added quality-metadata normalization for training records.
- Stopped all churn/SLM processes at end of session by user request.
Core changes this session
-
New corpus + orchestration:
datasets/run_specs/(base families)datasets/example_run_specs/(project-scoped specs)example_projects/(10 curated fixtures)tools/mcp/generate_run_specs.shtools/mcp/generate_example_projects_and_specs.shtools/mcp/run_spec_batch_with_capture.shtools/mcp/synthesize_bad_from_run_specs.sh
-
Pipeline behavior updates:
tools/mcp/run_sprint_taskitem_pipeline.shaccepts arbitrary markdown input files (not just sprint plans).tools/mcp/run_spec_batch_with_capture.shuses unique summary names (timestamp + pid + random) to prevent overwrite.tools/mcp/run_novelty_churn_loop.sh:- includes
datasets/example_run_specsfamilies (INCLUDE_EXAMPLE_PROJECT_SPECS=1) - emits UTC heartbeats with counters
- tolerates per-batch failures without exiting loop
- runs
reclassify_tool_execution_quality.sheach cycle
- includes
-
Data quality enrichment:
tools/mcp/reclassify_tool_execution_quality.shnow assigns:tool_call_success(bool)failure_class(schema|timeout|transport|policy|logic|null)recovery_pattern(retry|reformulate_args|choose_different_tool|abort|unknown|none)failure_terminality(recoverable|terminal|unknown|null)quality_schema_version=2
training_data/lora/README.mdupdated for new fields.
Current dataset snapshot (end of session)
training_data/lora/mcp_call_quality.jsonl: 85k+ rowstraining_data/lora/mcp_calls_bad.jsonl: 16k+ rows, ~2.9M bad-token estimate (char/4 heuristic)training_data/lora/mcp_multimodel_balanced.jsonl: category-balanced curated subset
Runtime observations
- Qwen 14B SLM remains mostly
final_status=blockedin recentslm_batch_*.jsonlmanifests. - Most bad-data volume currently comes from spec/synthetic families, not successful SLM tool execution.
- Infra-noise filter remained effective (
curl(7)-style contamination stayed near zero).
Tomorrow quick-start checklist
- Ensure Ollama daemon and model are available:
export OLLAMA_MODELS=/home/bill/.ollama/models
export OLLAMA_HOST=127.0.0.1:11435
export OLLAMA_MODEL='qwen2.5-coder:14b'
ollama ps || true
curl -sS --max-time 3 http://127.0.0.1:11435/api/tags | jq -r '.models[]?.name'
- Sanity-check runtime wiring:
cd /home/bill/Documents/CLionProjects/whetstone_DSL
OLLAMA_URL=http://127.0.0.1:11435/api/generate \
OLLAMA_TAGS_URL=http://127.0.0.1:11435/api/tags \
./tools/mcp/validate_slm_runtime.sh
- Start churn loop (foreground with live log tee):
cd /home/bill/Documents/CLionProjects/whetstone_DSL
OLLAMA_HOST=127.0.0.1:11435 \
OLLAMA_URL=http://127.0.0.1:11435/api/generate \
OLLAMA_TAGS_URL=http://127.0.0.1:11435/api/tags \
CYCLES=200 FAMILIES_PER_CYCLE=4 RUN_SLM=1 SYNTH_BAD=1 \
BAD_PER_CATEGORY=400 GOOD_TO_BAD_RATIO=1.0 BAD_MAX_PER_SIGNATURE=20 BAD_MIN_CATEGORY_COUNT=10 \
ENABLE_INFRA_NOISE_FILTER=1 INCLUDE_EXAMPLE_PROJECT_SPECS=1 HEARTBEAT_SECONDS=60 \
./tools/mcp/run_novelty_churn_loop.sh | tee -a /tmp/novelty_churn_live.log
- Watch heartbeat and progress:
tail -f /tmp/novelty_churn_live.log | rg 'HEARTBEAT|NOVELTY CYCLE .* DONE'
If immediate goal is only bad-token target completion
- Keep current spec/synthetic flow unchanged.
- Skip extra refactors and prioritize uninterrupted run time.
- Run
reclassify_tool_execution_quality.shat least once after any manual ingestion.