jeval v0.1.0 — measure what your classifier's confidence is worth
Install
uvx --from git+https://github.com/rlaope/jeval jeval demoOr with pip, straight from this release:
pip install https://github.com/rlaope/jeval/releases/download/v0.1.0/jeval-0.1.0-py3-none-any.whl
jeval demoWhat it does
Your classifier answers with a confidence. jeval measures what that confidence is actually worth
on your own data, and — from what a mistake costs you — derives the confidence line at which a
decision should be automated instead of handed to a human.
Everything is computed from labeled decision records on disk. No server, no database, no network
call, no account.
Using it from an agent
The whole setup is one sentence to an agent:
Install jeval (
uvx --from git+https://github.com/rlaope/jeval jeval --help), find where my
classifier's requests and responses are logged, adapt.jeval/ingest-map.yamlto that shape, run
jeval report, and show me the report file.
Machine-readable entry point: https://github.com/rlaope/jeval/blob/main/llms.txt
What you get
| Artifact | What it answers |
|---|---|
report.html — one self-contained file |
Are the confidences trustworthy, where do they break, and what does the current threshold cost? |
thresholds.yaml — read by your application |
The threshold to deploy, with a bootstrap interval, and whether splitting by segment pays |
labels.csv — a sheet to fill in |
The decisions whose labels would teach the tool the most |
calibration-*.yaml — optional |
A correction map your application applies, exported only when the gain is real |
Verified
488 tests, ruff + mypy clean, CI green on 3.10 / 3.11 / 3.12.