Releases: dave8172/doceval
Release list
v0.4.2 — first PyPI release
The first release on PyPI.
pip install docevalFixed
- The
[mcp]extra was broken on a fresh install.pyproject.tomldeclaredmcp>=1.0, so a new install resolved MCP 2.x — whereFastMCPwas renamed toMCPServer— whilemcp_server.pyis v1 code. Anyone runningpip install "doceval[mcp]"got anImportError. Now pinned tomcp>=1.0,<2. Migrating to the 2.x API is separate work. - LICENSE file added. The README and
pyproject.tomlboth said MIT with nothing behind them.
Changed
- Package metadata cleaned up ahead of the first upload, since published metadata cannot be edited afterwards: SPDX
license = "MIT"withlicense-files, anauthorsentry, Repository and Changelog URLs, Python version classifiers, and adescriptionthat matches what the harness does — it still read "document extraction pipelines". - Published from GitHub Actions over OIDC trusted publishing; no API token is stored anywhere.
What v0.4.0 changed, if you missed it
The input layer widened beyond documents. Emails, HTML pages, call transcripts and feed dumps now evaluate exactly like scanned invoices, because what doceval scores is whether the extracted fields are right — and that never depended on the file type. examples/leads/ is the worked non-document case.
v0.4.1 — first PyPI release
Same code as v0.4.0. This release exists to make pip install doceval true.
pip install docevalAdded
- LICENSE file. The README and
pyproject.tomlboth said MIT and nothing backed it. - PyPI packaging, published from GitHub Actions over OIDC trusted publishing — no API token stored anywhere.
Changed
- Package metadata cleaned up ahead of the first upload, since published metadata cannot be edited afterwards: SPDX
license = "MIT"withlicense-files(replacing the deprecated table form and theLicense :: OSI Approvedclassifier), anauthorsentry, Repository and Changelog URLs, Python version classifiers, and adescriptionthat matches what the harness actually does — it still read "document extraction pipelines".
What v0.4.0 changed, if you missed it
The input layer widened beyond documents. Emails, HTML pages, call transcripts and feed dumps now evaluate exactly like scanned invoices, because what doceval scores is whether the extracted fields are right — and that never depended on the file type. examples/leads/ is the worked non-document case.
v0.4.0 — evaluate any input, not just documents
doceval scores whether extracted fields are correct. That never depended on the source being a document — the scoring, failure-taxonomy and cost layers were always format-blind, and the only document-specific thing in the codebase was a seven-extension allowlist.
So it's gone.
Added
- Text inputs:
txt,md,json,jsonl,ndjson,csv,tsv,html,htm,xml,eml,msg,vtt,srt— alongside the existing PDF/PNG/JPG/TIFF/WEBP. An email, an HTML page, a call transcript or a feed dump now evaluates exactly like a scanned invoice. examples/leads/— the worked non-document case. Eight inbound sales emails: a forwarded thread where the lead is the original sender rather than the forwarder, a terse RFQ, a web-form dump, an automated tender notice with no contact details, and one enquiry in Italian. Three carry null-heavy labels on purpose, so an extractor that invents a quantity or a budget scores worse.
Fixed
- The CLI counted inputs with
iterdir()whilerun_evaldiscovers withrglob(), so the "N found" banner disagreed with what was actually evaluated whenever inputs sat in subdirectories. The CLI now imports the harness's extension set instead of duplicating it inline. - The README told you to
pip install doceval. It is not on PyPI. Install from the repo:
pip install git+https://github.com/dave8172/docevalUnchanged, deliberately
The name, and what this is for: field-level extraction correctness. The public document benchmarks score page-parsing fidelity — did you reproduce the layout, the table, the formula — and none of them score whether the extracted field is actually right. That is the gap doceval is pointed at, and widening the inputs doesn't widen the claim.