ctx v2.3.0
ctx v2.3.0 — Welle 48 Empirie komplett (V6-opt-in + Multi-Domain + Korpus-Sweep)
Minor release. Documentation + bench-pointers only. NO code/API changes
to ctx prod since v2.2.0. Bench results live in .project/bench-session-crag/.
Welle-48 results bundle:
W48-01 V6-Prompt: graded-confidence opt-in via CTX_PROMPT_VERSION=v6.
Movie n=10: +0.1pp over V5.2. PARTIALLY CONFIRMED.
W48-02 Multi-Domain CRAG: finance/sports/open n=10 each.
V6 vs V5.2 Δ averaged +0.075pp (+5 wins / -3 losses).
H6 (Domain-Bias): finance FALSIFIED (predicted strength, actual -0.20),
sports PARTIALLY-FALSIFIED, open CONFIRMED (V6 catches false-premise).
Key insight: Generator (qwen3.5:9b) is bottleneck, not Retrieval.
W48-03 Korpus-Sweep n=50/100/200:
Crossover-Point: RRF beats Cosine bereits bei n=50.
Δ (RRF-Cosine): +0.167 (n=50), +0.267 (n=100), +0.267 (n=200).
Counter-intuitive: Slot-Recall lower (RRF 55-75%, Cosine 69-82%),
but End-to-End score higher (RRF +0.167-0.267pp).
Mechanismus: V5.2 Refusal-Disposition profits from RRF diversification
under CRAG-Score (2c+m)/n-1 (Refusal=0 > Hallu=-1).
W48-04 Embedding-Pipeline-Audit: LOW severity. W47-NEU-D Hypothese
('Ollama full vs llama Q4_K_M Drift') FALSIFIZIERT. Both pipelines
use Q4_K_M, Δ-Cosine = 0.001-0.002 (Noise-Floor).
W48-05 Selective Re-Embed: NO-ACTION (W48-04 LOW severity).
Killer-Welle-49-Items derived:
- W49-A: n≥50 per domain for statistical significance
- W49-B: Generator-Eskalation qwen3.5:9b → qwen3.6:27b for numeric-extraction
- W49-C: V6.1 Generic-Examples (domain-agnostic)
- W49-D: Dream-Pick-Balance (round-robin for multi-domain)
- W49-E: False-Premise Detection Channel
- W49-X1: Cosine + V6 re-bench
- W49-X2: CTX_CONFIDENT_THRESHOLD sweep
- W49-X4: Dream-mature re-bench at n=200 (planned next)
No new migrations. No breaking changes.
Installation
With Go:
go install github.com/GottZ/ctx/cmd/ctx@v2.3.0Binary download:
Download the binary for your platform, make it executable, move to PATH:
chmod +x ctx-*
sudo mv ctx-* /usr/local/bin/ctx