Skip to content

jeval v0.1.0 — measure what your classifier's confidence is worth

Choose a tag to compare

@rlaope rlaope released this 21 Sep 00:15
· 85 commits to main since this release

Install

uvx --from git+https://github.com/rlaope/jeval jeval demo

Or with pip, straight from this release:

pip install https://github.com/rlaope/jeval/releases/download/v0.1.0/jeval-0.1.0-py3-none-any.whl
jeval demo

What it does

Your classifier answers with a confidence. jeval measures what that confidence is actually worth
on your own data, and — from what a mistake costs you — derives the confidence line at which a
decision should be automated instead of handed to a human.

Everything is computed from labeled decision records on disk. No server, no database, no network
call, no account.

Using it from an agent

The whole setup is one sentence to an agent:

Install jeval (uvx --from git+https://github.com/rlaope/jeval jeval --help), find where my
classifier's requests and responses are logged, adapt .jeval/ingest-map.yaml to that shape, run
jeval report, and show me the report file.

Machine-readable entry point: https://github.com/rlaope/jeval/blob/main/llms.txt

What you get

Artifact What it answers
report.html — one self-contained file Are the confidences trustworthy, where do they break, and what does the current threshold cost?
thresholds.yaml — read by your application The threshold to deploy, with a bootstrap interval, and whether splitting by segment pays
labels.csv — a sheet to fill in The decisions whose labels would teach the tool the most
calibration-*.yaml — optional A correction map your application applies, exported only when the gain is real

Verified

488 tests, ruff + mypy clean, CI green on 3.10 / 3.11 / 3.12.