Skip to content

Releases: dave8172/doceval

v0.4.2 — first PyPI release

Choose a tag to compare

@dave8172 dave8172 released this 04 Sep 12:24

The first release on PyPI.

pip install doceval

Fixed

  • The [mcp] extra was broken on a fresh install. pyproject.toml declared mcp>=1.0, so a new install resolved MCP 2.x — where FastMCP was renamed to MCPServer — while mcp_server.py is v1 code. Anyone running pip install "doceval[mcp]" got an ImportError. Now pinned to mcp>=1.0,<2. Migrating to the 2.x API is separate work.
  • LICENSE file added. The README and pyproject.toml both said MIT with nothing behind them.

Changed

  • Package metadata cleaned up ahead of the first upload, since published metadata cannot be edited afterwards: SPDX license = "MIT" with license-files, an authors entry, Repository and Changelog URLs, Python version classifiers, and a description that matches what the harness does — it still read "document extraction pipelines".
  • Published from GitHub Actions over OIDC trusted publishing; no API token is stored anywhere.

What v0.4.0 changed, if you missed it

The input layer widened beyond documents. Emails, HTML pages, call transcripts and feed dumps now evaluate exactly like scanned invoices, because what doceval scores is whether the extracted fields are right — and that never depended on the file type. examples/leads/ is the worked non-document case.

v0.4.1 — first PyPI release

Choose a tag to compare

@dave8172 dave8172 released this 04 Sep 12:22

Same code as v0.4.0. This release exists to make pip install doceval true.

pip install doceval

Added

  • LICENSE file. The README and pyproject.toml both said MIT and nothing backed it.
  • PyPI packaging, published from GitHub Actions over OIDC trusted publishing — no API token stored anywhere.

Changed

  • Package metadata cleaned up ahead of the first upload, since published metadata cannot be edited afterwards: SPDX license = "MIT" with license-files (replacing the deprecated table form and the License :: OSI Approved classifier), an authors entry, Repository and Changelog URLs, Python version classifiers, and a description that matches what the harness actually does — it still read "document extraction pipelines".

What v0.4.0 changed, if you missed it

The input layer widened beyond documents. Emails, HTML pages, call transcripts and feed dumps now evaluate exactly like scanned invoices, because what doceval scores is whether the extracted fields are right — and that never depended on the file type. examples/leads/ is the worked non-document case.

v0.4.0 — evaluate any input, not just documents

Choose a tag to compare

@dave8172 dave8172 released this 04 Sep 11:08

doceval scores whether extracted fields are correct. That never depended on the source being a document — the scoring, failure-taxonomy and cost layers were always format-blind, and the only document-specific thing in the codebase was a seven-extension allowlist.

So it's gone.

Added

  • Text inputs: txt, md, json, jsonl, ndjson, csv, tsv, html, htm, xml, eml, msg, vtt, srt — alongside the existing PDF/PNG/JPG/TIFF/WEBP. An email, an HTML page, a call transcript or a feed dump now evaluates exactly like a scanned invoice.
  • examples/leads/ — the worked non-document case. Eight inbound sales emails: a forwarded thread where the lead is the original sender rather than the forwarder, a terse RFQ, a web-form dump, an automated tender notice with no contact details, and one enquiry in Italian. Three carry null-heavy labels on purpose, so an extractor that invents a quantity or a budget scores worse.

Fixed

  • The CLI counted inputs with iterdir() while run_eval discovers with rglob(), so the "N found" banner disagreed with what was actually evaluated whenever inputs sat in subdirectories. The CLI now imports the harness's extension set instead of duplicating it inline.
  • The README told you to pip install doceval. It is not on PyPI. Install from the repo:
pip install git+https://github.com/dave8172/doceval

Unchanged, deliberately

The name, and what this is for: field-level extraction correctness. The public document benchmarks score page-parsing fidelity — did you reproduce the layout, the table, the formula — and none of them score whether the extracted field is actually right. That is the gap doceval is pointed at, and widening the inputs doesn't widen the claim.