What's Changed
Added
- Training web API to list ongoing trainings and interrupt a running one, releasing the per-model lock #1483
- Development/model debug API exposing the first-level models for inspecting intermediate results #1439
- Push-based export of JVM/process runtime metrics over OTLP, complementing the Prometheus scrape endpoint #1479
/metrics/prometheusendpoint now serves the Prometheus exposition format with JVM/process metrics, plus Prometheus/Grafana setup docs #1473- Linking-aware
affiliation_linkedmetric in the end-to-end HEADER evaluation #1493 #1467 - Rootless Docker images (Kubernetes/OpenShift friendly) #1442
- Fail fast at startup when the grobid-home path contains spaces and Wapiti is the configured engine #1481
- Apache 2.0 licence headers in source files #1485
- Documentation: community page #1447, PDF-TEI Editor reference #1448, and instructions for adding new model flavors #1465
- CI: CodeQL analysis workflow #1475, dependabot updates for GitHub Actions #1494, and runs on forks without publishing credentials #1451
Changed
- Updated pdfalto to 0.6.2. Notable for GROBID: deterministic output (an uninitialised read and a use-after-free made line/block grouping depend on heap contents, so the same PDF could yield different
<TextLine>/<TextBlock>structure across runs), bounded peak memory on very large or vector-heavy documents, superscript citation/affiliation callouts separated from adjacent words (token boundary now driven by pdfalto's own superscript detection), and several text-extraction fixes. The command-line options GROBID passes are unchanged. - pdfalto is now given GROBID's configured temp directory (
grobid.temp, by defaultgrobid-home/tmp) throughTMPDIR. Since 0.6.1 pdfalto streams large page DOMs to a scratch file to bound peak memory and picks its location fromTMPDIR, falling back to/tmp— which is tmpfs in most containers, so spilling there leaves peak memory unchanged and risks ENOSPC. - Reworked author–affiliation linking into a dedicated, unit-tested
AuthorAffiliationAssignerwith a priority-based strategy #1467 - Preserve extracted affiliations when header consolidation rewrites authors, via staged reconciliation #1488
- Parallelized the end-to-end evaluation scoring phase (~20 min → ~2 min), with upfront setup verification in the Python script #1487
- Rewrote sentence-segmentation re-alignment as a drift-free forward two-pointer alignment and made language detection thread-safe #1457
- Language-sensitive tokenization in
PDFALTOSaxHandlerinstead of always using the default analyzer #1480 - Only one training per model at a time; a concurrent request for the same model returns 409 Conflict #1477
- Unified training-data file generation between the web API and batch mode #1508
- Lexicon: added
Lexicon.builder()for optional eager gazetteer pre-loading (lazy stays the default);getInstance()is now deprecated #1440 - Lexicon: added 4 missing ISO 3166-1 country codes (BQ, CW, SS, SX) and migrated AN to its ISO 3166-3 form ANHH
- Documentation: expanded the end-to-end evaluation and configuration guides and recorded the
grobid-evaluationdataset DOI #1419 #1430 #1501 - Dependency updates: OpenNLP 1.9.4 → 2.5, Jetty, jackson 2.21.4, DeLFT 0.4.6; dropped unused dependencies #1449 #1469 #1423
- Replaced Powermock with Mockito #1458
- Adopted Spotless for code formatting #1384, rewrote overloaded methods #1401, and added codespell spell-checking #1365 #1450 #1478
Fixed
- Handle the exit codes pdfalto 0.6.1 introduced. Code 5 reports that the ALTO was written correctly but page streaming was disabled mid-run, so peak memory was no longer bounded; both execution paths rejected any non-zero exit and would have discarded these successful conversions. It is now a logged warning. Codes 4 (ALTO write failed) and 98 (allocation failure) are reported as
PDFALTO_CONVERSION_FAILUREinstead of falling through toBAD_INPUT_DATA, which blamed the PDF for a pdfalto-side failure. - JVM shutdown deadlock when closing JEP Python interpreters: close now runs on the owning worker thread and is idempotent #1506
- HTTP 500 from
processFulltextDocumentwhen two footnotes share the same superscript marker; such callouts now fall back to plain text #1472 - TEI paragraph boundaries no longer collapse when a paragraph starts right after a trailing reference marker #1482
- Invalid/unbalanced XML in generated training data across models #1470, and unclosed
<bibl>before<other>in reference-segmenter training data #1466 xml:idvalues in generated training files are now valid NCNames #1508- NPE in reference-segmenter training-data generation after segmentation retraining #1490
- Document language is detected for segmentation training data instead of hardcoding
xml:lang="en"#1460 - Textual year/month/day fields stay in sync with the normalized publication date for library callers and non-TEI output paths #1463
TextUtilities.clean()no longer folds the letters æ/Æ and œ/Œ to ASCII "ae"/"oe"; typographic ligature expansion (fi/fl/ff) is unchanged #1461- Guard against out-of-range page index in citation annotation, which surfaced as HTTP 500 on malformed PDFs #1459
- Model selection with mixed DeLFT/Wapiti engines and flavor selections, with clearer logging when a flavor falls back to the base model #1455
- Citations consisting only of non-breaking spaces now return 204 No Content instead of HTTP 500 #1407
- biblio-glutton health probe in the evaluation configuration check now targets an existing endpoint (
/service/data) instead of always reporting a healthy glutton as unreachable #1492 - Block/segmentation desync warning now includes the page number and a text excerpt so occurrences can be located and reproduced #1471
./gradlew install#1427, git revision information #1433, and Docker image summary #1429
Security
- Prevent command injection through crafted PDF file names in the non-server
pdfaltopath: the command is no longer interpolated into abash -cstring but passed as positional parameters and exec'd via"$@"(GHSA-mgxf-7mg7-qpmf) #1477 - Stop leaking a JVM thread per request on the
/api/modelTrainingendpoint by shutting down the per-request executor (GHSA-g2r5-4c8r-c84f) #1477 - Remove the vulnerable JLine telnet server module from the classpath by depending on
jline-terminalonly instead of theorg.jline:jlineuber-jar pulled in transitively byprogressbar(GHSA-47qp-hqvx-6r3f, GHSA-2r2c-cx56-8933) #1469 - Upgrade jackson (core, databind, afterburner, dataformat-yaml) to 2.21.4 to address CVE-2026-54513 (array-element type allowlist bypass in polymorphic type validation) #1469
- Upgrade Apache OpenNLP to 2.5 (arbitrary class loading via model manifest) and Jetty (HTTP request smuggling via chunked extension quoted-string parsing) #1449
- Harden
ZipUtilsagainst zip-slip: each entry's canonical output path is validated against the target directory before any write (flagged by CodeQL) #1486
New Contributors
- @flrjrf made their first contribution in #1427
- @luismmontilla made their first contribution in #1430
- @yarikoptic made their first contribution in #1365
- @lfoppiano with @Copilot made their first contribution in #1480
Full Changelog: 0.9.0...0.9.1