Skip to content

Releases: Alberto-Codes/typevet

v0.7.0

Choose a tag to compare

@Alberto-Codes Alberto-Codes released this 01 Oct 13:43
ee3e720

0.7.0 (2026-10-01)

Features

  • adapters: expose Retry-After and rate-limit headers on a vLLM HTTP error (8a273df), closes #355
  • adapters: record the request id typevet sends in receipts and errors (68aab28), closes #356
  • adapters: refuse a Gemma 4 framing that lacks the no-thinking prefill (8edd342), closes #354
  • domain: make the option block and the context template substitutable named parts (2db4a5e), closes #364 #360
  • evals: evolve criteria beside instructions through a multi-part wording candidate (112bca9), closes #363
  • evals: pin seed and evolved wording by digest in wording receipts (ea38912), closes #362 #360
  • ports: carry off_option_threshold through the port protocol and every wrapper (3df3ff7), closes #368 #353
  • runtime: apply a caller-supplied calibration map to Noul answers (f1a6d6c), closes #352
  • runtime: apply a Score calibration map and write map artifacts from receipts (00f7725), closes #352
  • scoring: flag an answer whose off-option mass exceeds a caller threshold (722728b), closes #353

Fixes

  • evals: mask an echoed header name used as a JSON key in the acceptance receipt (bbd1d19), closes #351

Documentation

  • reference: list the text parts of a judgment call and their evolution status (f55a9c0), closes #361 #360

v0.6.0

Choose a tag to compare

@Alberto-Codes Alberto-Codes released this 30 Sep 21:53
53034de

0.6.0 (2026-09-30)

Features

  • adapters: reach a vLLM server behind an API gateway (58eebce), closes #331
  • evals: post-hoc calibration of the committed Noul receipts (78aea74), closes #343
  • evals: record the vLLM server configuration in receipts (d34c3b9), closes #341
  • evals: seed the synthetic check run from the environment (1dfbc36), closes #344

Fixes

  • adapters: close the owned scoring client under the lock (5f5caaa), closes #345
  • adapters: create the lazy scoring client under a lock (6b44d72), closes #340
  • adapters: mask gateway header values as whole tokens and protect owned headers (2719ad6), closes #348
  • evals: one seed rule for the check sheets command and the live run (6fdc464), closes #349
  • evals: refuse a wrong split and a two-engine gauge in receipts (e030ee4), closes #339

Documentation

  • explanation: record the #342 design outcome on the throughput page (9c6afa5)
  • reference: receipt blocks page, and a timestamped metrics sample (2492d95), closes #347

v0.5.0

Choose a tag to compare

@Alberto-Codes Alberto-Codes released this 30 Sep 19:59
85bc554

0.5.0 (2026-09-30)

Features

  • evals: concurrency option for the face, check and signature runners (b287486), closes #334
  • evals: Jev judge option for the wording evolution live test (d86e70c), closes #328
  • evals: Jev vs Gemma held-out comparison harness with Cohen's kappa (d86aa9f), closes #329
  • evals: served-template probe and native Gemma framing for text wording runs (0628eb0), closes #327
  • evals: throughput fields in the image receipts (93a13fb), closes #335

Documentation

  • explanation: Gemma 4 and Jev on DIFrauD, each with its evolved wording (44ffb4d), closes #330 #252 #324
  • explanation: image judgment throughput on one H100 (50834b9), closes #337 #332
  • explanation: latency and cost of the 0.4.0 experiments (4f66fa0), closes #326

v0.4.0

Choose a tag to compare

@Alberto-Codes Alberto-Codes released this 30 Sep 16:01
f1292e7

0.4.0 (2026-09-30)

Features

  • evals: CEDAR signature pair loader and two-image request builder (46de94e), closes #318
  • evals: check-versus-register runner, metrics and opt-in live test (38cada3), closes #316
  • evals: DIFrauD train, validation and held-out splits without #236 overlap (d14dcf0), closes #307
  • evals: gepa-adk wording evolution runner with Brier scorer and length cap (6f9f740), closes #308
  • evals: harness for the DIFrauD wording evolution and its held-out check (502ebc9), closes #309
  • evals: signature-pair runner, metrics and opt-in live test for CEDAR (e3e4643), closes #319
  • evals: synthetic check generator and check-versus-register request builder (b3d5457), closes #315
  • evals: top-label and classwise ECE for Choice and Score above 10 options (38f9272), closes #296
  • scoring: report off-option probability mass in scoring results (db5ff12), closes #297

Fixes

  • adapters: exact vocabulary size for the llama.cpp off-option completeness check (81a2e02), closes #321
  • adapters: media-marker refresh edge cases after a model swap (ce63ddd), closes #323
  • adapters: refresh the llama.cpp media marker after a router model reload (e3dd4cb), closes #322

Documentation

  • explanation: check images against a synthetic register and what the typed answers mean (7299fd6), closes #317 #316 #303
  • explanation: DIFrauD wording evolution and its pre-registered held-out check (3133817), closes #309
  • explanation: two-image signature comparison on CEDAR and what its probability means (c418b0c), closes #320 #319 #304
  • reference: record the Molmo2-4B single-token control check (643b68a), closes #299

v0.3.0

Choose a tag to compare

@Alberto-Codes Alberto-Codes released this 30 Sep 11:54
c60236c

0.3.0 (2026-09-30)

Features

  • evals: LFW View 2 pair loader and two-image face-match requests (dc3d7ca), closes #300 #292

Fixes

  • adapters: map errors from the vision factory's /apply-template probe (64814a0), closes #311
  • adapters: map vision tokenizer errors and characterize runtime limits (597ce88), closes #204 #298
  • adapters: retry one early close on llama.cpp scoring calls (fdae661), closes #305 #310
  • evals: wider PSAI leak check, key-free receipt frames, prompt-read controls (09f6787), closes #295

Documentation

  • explanation: two-image face matching and what its probability means (cd76f82), closes #302 #292
  • integrations: explain why the judgevet bridge keeps text-only instructions (f7560b3), closes #291
  • split long sentences in the docs index, CORD smoke and judgevet pages (4ce5971), closes #295
  • state the 0.2.0 release where pages said 0.1.0 is current (04e8b81), closes #294
  • vllm: record the live KV-cache gauge reading on the tested pin (02db15d), closes #231

v0.2.0

Choose a tag to compare

@Alberto-Codes Alberto-Codes released this 30 Sep 04:03
fbc16e4

0.2.0 (2026-09-30)

Features

  • domain: native Choice up to 24 options with digit-then-letter controls (34f23ad), closes #286 #287 #237
  • integrations: a judgevet provider bridge over JudgmentPort (2b04383), closes #284
  • integrations: judgevet 0.15 conformance kit and per-question images (67728f5), closes #290

Fixes

  • adapters: drop the cause chain from masked vLLM errors when a key is set (71ee1fb), closes #276
  • adapters: reject non-finite vLLM timeouts and mask keys in set and bytes values (e8ecc20), closes #227
  • domain: validate CandidateScoringRequest.media items (ce3c9e6), closes #222
  • evals: harden the PSAI leak check, KV-cache reading and runner failure counting (dec6dc7), closes #223 #225 #238 #261
  • site: focus ring, Noul search rank, 2x header reflow and Home card headings (5bbfecb), closes #282 #267
  • site: keep the closed search panel out of the Tab order (3a7a67c), closes #285 #267 #242
  • site: keyboard access to the mobile navigation drawer (918d8d9), closes #283 #267
  • tests: keep API keys out of pytest failure output (599de01), closes #251

Documentation

  • correct the local alias identity, framing prefill and Gemma 4 image cost (867a61a), closes #233 #235 #260
  • explanation: correct the stale #129 template claim in native typed judgments (698d53d), closes #278
  • explanation: gather limits and known gaps on one page with page status (cb51e80), closes #247
  • index: list every page in docs/README.md and drop the stale PyPI claim (638a2e1), closes #275
  • reference: map vLLM adapter failures to typevet errors (367051e), closes #273
  • reference: one page for data handling, key masking and security (fa07066), closes #246
  • reference: one page for releases, changelog and versioning (40f7286), closes #274
  • reference: split the Python API reference into one page per package (f7af6d4), closes #266
  • site: define accessible light and dark visual tokens (c7dceae), closes #263 #262
  • site: give the Home routes a primary start action and task cards (acc7e25), closes #264
  • site: group reference navigation around reader tasks (219bf76), closes #265
  • tutorial: first typed judgment on a local llama.cpp server (0c1b529), closes #281
  • tutorial: make the offline tutorial a full first run for a new reader (0ac1f92), closes #280

v0.1.0

Choose a tag to compare

@Alberto-Codes Alberto-Codes released this 29 Sep 23:46
2ecfc99

typevet 0.1.0

The first full release of typevet on PyPI. typevet is pre-1.0, so the import surface can change in a minor release.

typevet asks a model typed questions and returns typed answers:

  • Noul (yes or no), Choice (one label) and Score (one rubric level). typevet computes each answer from the model's next-token probabilities, before sampling.
  • JSON objects that pass a JSON Schema that you supply. If the output does not pass, typevet raises an error.

Install

pip install typevet

typevet requires Python 3.12 or later. CI tests Python 3.12.

Try it without a model

This example uses a scripted fake instead of a model server:

from typevet.domain import Noul
from typevet.runtime import ScoringJudgmentAdapter
from typevet.testing import ScriptedScoringFake

fake = ScriptedScoringFake(logprobs={"True": -0.2, "False": -1.0})
port = ScoringJudgmentAdapter(fake, tokenize_content=lambda t: (ord(t[0]),))
question = {"q": Noul(instructions="Ok?", criteria={"true": "Y", "false": "N"})}
result = port.judge("text", question, "fake")
print(result.nouls["q"].noul)  # probability of yes, near 0.69

Before you connect a model

typevet calls a model server that you run. It does not start one.

  • llama.cpp is the default backend. Set TYPEVET_LLAMA__BASE_URL if the server is not at http://127.0.0.1:8090.
  • For vLLM, set TYPEVET_BACKEND=vllm and TYPEVET_VLLM__BASE_URL. vLLM has no default URL.

Configuration lists every setting.

Tested pins

Backend Tested pin
vLLM vllm/vllm-openai:v0.30.0, BF16 google/gemma-4-31B-it, one H100 80 GB
llama.cpp, generation Build b11243, google/gemma-4-31B-it-qat-q4_0-gguf Q4_0, one A40 48 GB
llama.cpp, image input Build b11223, gemma-4-31b-kv9-q4km-mm

Each pin has one receipt. The support matrix gives the full pins and their limits.

Performance

On one H100 with 64 requests in flight, typevet judged 480 Banking77 records in 12.1 s (39.6 records/s), with 0 errors. Banking77 calibration passed. DIFrauD SMS calibration failed parity (ECE 0.158 against a threshold of 0.10). The results come from one run on one pod with one pin. See Performance on one H100.

Import surface

  • The root package exports nine names and __version__. The question types Noul, Choice and Score are in typevet.domain.
  • Supported imports lists every supported package and name. A contract test keeps that page equal to the __all__ of each package.
  • The wheel holds only the library: domain, ports, adapters, runtime and testing. The evaluation harnesses and dataset loaders are in the private typevet-evals workspace member, which is not published.
  • ADR 0002 records the package layout.

What is verified, and what is not

  • A valid structure does not prove accuracy or calibration. The receipts are small samples.
  • Other GPUs, models, precisions and server versions are not tested.
  • 0.1.0.dev1 on PyPI reserved the name. It is an early snapshot. Do not use it.

Documentation: https://alberto-codes.github.io/typevet/