Skip to content

Releases: belalik/erga

erga 0.8.0

Choose a tag to compare

@belalik belalik released this 26 Sep 08:55

This release finds the works a profile fetch cannot see, and checks curation files when they load rather than when a record happens to match.

  • verify lists works that name a member but link no author. OpenAlex sometimes prints a name in a byline without attaching it to any author profile, so no fetch by profile reaches the work. Once a build has written publications.json, verify searches bylines for each configured name and alias, runs the finds through the build's own curation (your overrides' exclusions and keep_distinct, exclude_types), drops whatever the list already holds by id, DOI or title, and lists the rest for your manual file. It also names a listed copy that lacks a DOI when a record carrying one exists, and a listed work that does not credit the member because its byline prints the name another way. The section is advisory, since a namesake's byline looks the same, and a name too common to judge (over 1,000 matching works) is skipped with a count. Write names with their diacritics, as bylines print them: the search also tries each name without them, but cannot restore accents a configured name lacks.
  • Curation files are checked when they load. Manual entries and overrides go through one check per field at load, so a typo under a stale id no longer passes a green build. A quoted date must be a calendar date (2025-02-31 is refused), year: true is refused in both files, and a year patched on its own must agree with the fetched date (patch date too, or set date: null). A manual entry with only a date takes its year from it. A file that built under 0.7.0 can be refused now; the message names the entry and the field.
  • An authors list written as one comma-joined string warns, in manual entries and overrides alike, since it tracks nobody. Two commas say "may be several authors", since a suffix such as "Jr." reads the same.
  • A blank alias is refused in erga.yml, as is one with no name in it.
  • Only the DOI is kept when OpenAlex glues HTML onto the field, as it does on some repository records, so DOI dedup, overrides, tags and the Crossref backfill all see it. The build says how many works needed it.
  • publications.json gets the permissions a plain write would give: an existing file keeps its mode and a new one follows the umask, instead of the owner-only mode a web server running as another user could not read.

Python ≥ 3.10. Install: pip install erga or uv tool install erga, or use the Action: uses: belalik/erga@v0.8.0 with version: "0.8.0".

erga 0.7.0

Choose a tag to compare

@belalik belalik released this 23 Sep 07:07

This release lets a declared home follow a career that moved, so someone who joined recently no longer reads as somebody else's profile.

  • home: takes a list of ROR ids as well as one, at the top level or per author: home: [<your ROR>, <their previous ROR>]. Every listed place is home, each with its whole country, as a single declaration always was. Bare ids and ror.org URLs mix freely, and one id named twice counts once. An empty list is refused, since null already says "no declaration". A per-author list replaces the top-level default, so an arrival's list repeats the department.
  • The wrong-profile warning names the likeliest reading first: the declaration is incomplete (someone who moved here needs their previous institutions listed under home:), before a wrong declaration or a profile that is not only theirs. It names the declared home countries it measured against. With one place declared, every verdict is the same as in v0.5.0; only this wording changed.
  • Corrected: home: null does not remove the check. It returns that author to the inferred rule, as it has since v0.5.0; the design record said otherwise.

Python ≥ 3.10. Install: pip install erga or uv tool install erga, or use the Action: uses: belalik/erga@v0.7.0 with version: "0.7.0".

erga 0.6.0

Choose a tag to compare

@belalik belalik released this 22 Sep 12:57

This release turns the weekly refresh into something a reviewer can read: every build says what changed since the output it found in place.

  • A headline on every build (since previous output: 3 added, 1 removed, …), and build --summary PATH writes the full Markdown page: the run's warnings first, then additions, removals, retractions and tracking changes listed one by one with their audit tell (tracked person, co-authors, venue, DOI, a manual-entry mark), then every other change OpenAlex made counted per field as drift, then tag changes as the maintainer's own. Listed sections cap at 150 entries. A shrink warning, never a refusal, when more than a tenth of the previous works are gone, since a degraded fetch reads as mass removals.
  • erga diff OLD NEW: the same page for any two output files, with no config or network.
  • The Action's summary input: a path to write the page, so a pull-request workflow can use it as the PR body (README, recipe 2).
  • An unreadable publications.json in place is reported, not passed off as a first build. build warns and goes on; erga diff treats it as an error. A file whose works or bylines are not the shapes erga writes counts as unreadable rather than being partly read.

Python ≥ 3.10. Install: pip install erga or uv tool install erga, or use the Action: uses: belalik/erga@v0.6.0 with version: "0.6.0".

erga 0.5.0

Choose a tag to compare

@belalik belalik released this 21 Sep 12:36

This release lets a config say where its people work, so the contamination check reasons from a declared home instead of a counted one.

  • home:, a ROR id, at the top level or per author (home: null opts one person out). The work-level contamination check then knows where home is instead of inferring it by majority, and a profile whose placed works are mostly elsewhere is reported as a wrong profile, with no exclusion advice, rather than having its majority listed as strangers. The declaration is read only by this check; like the ORCID it is trusted as given and is never evidence that a profile is the right person. A ROR that names no OpenAlex institution, or one without a country, aborts the build.
  • Measured before shipping on two pinned cohorts of forty random authors each: a declaration changed no verdict a career already had and added no cluster, and declaring a wrong home produced the wrong-profile verdict on 53 of 56 careers with enough evidence. The other thing it was built to do, lifting the check's silence on a scattered career, did not come up once in eighty random careers, so that remains unobserved rather than shown. Thin records stay silent whoever declares them.

Python ≥ 3.10. Install: pip install erga or uv tool install erga, or use the Action: uses: belalik/erga@v0.5.0 with version: "0.5.0".

erga 0.4.0

Choose a tag to compare

@belalik belalik released this 02 Sep 09:40

This release hardens intake at department scale, where contaminated ORCIDs and same-name profiles are a statistical certainty rather than an anomaly, and adds a check for the one failure verify cannot see.

  • Work-level contamination check, advisory: build now warns about clusters of works tied to one institution that share no collaborator and no institution with the rest of a profile. That is what a same-name stranger's works look like inside a correctly identified profile, which verify cannot see because the names match. The rule was settled on the one career known to carry contamination and re-measured for noise at about a quarter of a cluster per author, an upper bound; the check stays silent where the career is not the clear majority of the profile. Judging the cluster stays with the maintainer, by DOI in the overrides file.
  • verify tells a split profile from an iD carried by strangers by comparing profile names against the configured name and aliases, and name-searches every configured author for same-name profiles the config does not cover.
  • output.exclude_types drops noise types wholesale; manual entries and explicit exclude: false overrides are exempt.
  • Quieter builds: the raw type other no longer warns on every build.

The README now says where erga's job stops: finding an iD is the consumer's step, an ORCID is trusted as given, and the discovery procedure that worked for a department with iDs on file for five of sixty-three staff is written down.

Python ≥ 3.10. Install: pip install erga or uv tool install erga, or use the Action: uses: belalik/erga@v0.4.0 with version: "0.4.0".

erga 0.3.0

Choose a tag to compare

@belalik belalik released this 09 Aug 13:17

This release packages erga as a composite GitHub Action, so a site's CI can build its publications list without installing anything by hand.

  • GitHub Action: uses: belalik/erga@v0.3.0 runs erga build through uvx. The version input pins the erga release and is deliberately required, since the Action tag and the erga version move independently; config points at your erga.yml, and an optional api-key passes an OpenAlex key. Recipes (including a Jekyll setup proven on the first consumer) are in docs/action.md.
  • Override matching fixed for merged records: an override carrying keep_distinct together with a patch recorded its match before dedup, so when its record lost a DOI merge the patch was silently dropped and the override reported as redundant. Matches are now settled after dedup; regression-tested.
  • Config samples use an unassignable ORCID (9999 prefix). ORCID's fictitious-researcher iD and the all-zeros iD are both carried by real OpenAlex profiles, so the previous sample fetched a stranger's works instead of failing.

The Action is exercised by this repo's own CI: a deterministic fixture-only job on every push, plus a weekly live job that checks OpenAlex still answers in the shape erga expects.

Python ≥ 3.10. Install: pip install erga or uv tool install erga, or use the Action.

erga 0.2.0

Choose a tag to compare

@belalik belalik released this 08 Aug 14:28

First release shaped by consumer feedback: the origin lab site completed its switch to erga (the v0.2 milestone), and both changes here come out of that adoption.

  • authors[].tracked_as: tracked authorships now carry the canonical configured author name alongside the tracked flag. erga holds the alias table and does the matching, so consumers can build reliable per-author filters without re-implementing alias logic against display-name variants. Additive schema field.
  • Reconstructed abstracts are now clean plaintext: HTML entities decoded (including double-encoded ones seen in the wild) and markup tags stripped, matching what publisher-supplied abstracts actually need.

Validated against the first consumer's full 187-work corpus: abstracts fully converge with its previous pipeline's output; every schema field compared.

Python ≥ 3.10. Install: pip install erga or uv tool install erga.

v0.1.0

Choose a tag to compare

@belalik belalik released this 05 Aug 17:00

First release. erga keeps a website's academic publications list current: authors (ORCID iDs) in one config file, works fetched from OpenAlex, normalized and deduplicated across registrars, curation applied from files that survive every automated refresh, and a canonical, diffable publications.json out.

  • erga build: config in, curated deterministic JSON out; --dry-run prints a summary without writing
  • erga verify: author-disambiguation report (split profiles, zero-work authors, implausible counts)
  • Curation files (manual.yml, overrides.yml, tags.yml) re-applied on every run
  • Crossref venue backfill; API etiquette built in (keys, delays, retries)
  • Validated against a production lab site's pipeline: 187/187 records converged, zero field diffs

Python ≥ 3.10. Install: pip install erga or uv tool install erga.