v4.7.0 — Calibration you can trust
Every confidence-like number the framework emits now carries its nature, and
the numbers that only looked like probabilities are gone. The reference is
Laya/Jev — typed decisions, proper scoring rules, calibration against
outcomes — adapted into a local-first, deterministic framework. No Kaggle
fine-tune, no external API in the runtime.
Changed
-
MetaCognition.calibration_score()no longer fabricates a prior. The
old1 - abs(E[C] - E[Y])macro-distance called a confident-wrong agent
"perfectly calibrated" when mean confidence equaled accuracy. It is now a
compatibility projection:NonebelowMIN_CALIBRATION_SAMPLES, otherwise
1 - ECE(5 equal-width bins).calibration()returns the full
ConfidenceValue(category, value, samples, metric). -
Confidence is multi-categoría.
ConfidenceValue(new
conscio/calibration.py) freezes four tiers:none(absence — value is
None, never a fabricated prior),asserted(declared deterministic
heuristic),derived(posterior over observed data),measured(ECE/Brier
against ground truth, samples + metric required).as_gate_input()raises
onnone— branch on the category, do not guess. ECE/Brier carry
lower_is_better=True;accuracydoes not. -
The council separates agreement from recommendation.
agreementis
1 - normalized_entropy(vote_counts)— four unanimous vetoes now read as
agreement 1.0 with avetorecommendation (the old table called it 0.1,
"disagreement").consensus_strengthsurvives as a deprecated alias. -
Coherence cold start is honest. The confidence field carries
ConfidenceValue—nonewhen evidence is insufficient. The0.85
default survives only in the legacy scalar projection that no gate reads;
theunmeasuredmechanism is preserved.
Fixed
-
The act fast-path no longer launders global calibration into per-action
safety. A globally "perfect" agent used to auto-pass tools it had never
run. Safety now comes from a per-tool Beta(1,1) posterior over the tool's
own ledger outcomes (p = (1+successes)/(2+attempts), derived, samples
exposed); zero attempts isnoneand never auto-executes.
AuditVerdict.confidenceisfloat | None— absence carries no number. -
Cold-start consumers of
calibration_score()no longer crash or guess.
TrustMatrix.autonomy_level(None >= 0.6 TypeError),max_action_retries
(int * None) andfast_path_ok()now treat absence as no-earned-trust:
autonomy stays L1, retries keep the warmup floor, the bypass gate fails. -
bench sabotage calibration treats
Noneconfidence as full suspicion.
Added
-
Outcome store (P1 ground truth): append-only
decision_outcomeswith
provenance —source(council/evaluate/squad/coherence),decision_ref
(unique, idempotent), immutablesnapshot,outcome
(pending/success/failure/reverted/false_positive),outcome_ts,
evidence_ref. The engine wires it and the council captures every decision
best-effort; verdicts arrive viaresolve()when the real outcome is known.
Pending is never a failure; a ghost resolve is visible, not silent.
Temperature-refit (Laya-styleT(task_type, option_count)) is deliberately
deferred until this store accumulates held-out outcomes. -
Vector-space signature:
{backend, model, dimension, version}persisted
on first write and validated before ingest/query — mixed-model signatures
are rejected before they corrupt recall.
Changed (embeddings)
- Embeddings are native-only by default.
EmbeddingProviderno longer
probes Ollama/LM Studio on boot: unsetCONSCIO_EMBED_BACKENDmeans
sentence_transformersin-process, zero network probes, and an explicit
failure when native is unavailable — never a silent daemon takeover.
CONSCIO_EMBED_BACKEND=ollama|openaiopts into a daemon;autois the
legacy fallback chain with a WARNING naming the selected backend.
semantic.pyuses the same factory instead of instantiatingOllamaEmbedder
directly.