v0.3.25
·
42 commits
to clean-main
since this release
feat: benchmark v3 stability + skills triggers in system prompt Benchmark v3 improvements: - Multi-run averaging (BENCHMARK_RUNS=3) with median aggregation to smooth LLM non-determinism - Relaxed thresholds (50-150% tolerance vs old 10-20%) to account for natural LLM variance - BENCHMARK_DISABLE=1 env var to skip checks entirely - BENCHMARK_STRICT=1 to enable blocking mode (default: warn-only, exit 0) - New npm script: benchmark:strict Skills loading: - Include trigger keywords in skill header output so agent can self-match at runtime - buildFullContextPrompt now supports task description for auto-matching - Updated agent prompt template to reference "Triggers" field Version: 0.3.25