Keeps a publications CV on Overleaf in sync with Google Scholar, from one command. Built to be re-run for years with as little manual work as possible: every step is idempotent, anything it cannot decide lands in WORKLIST.md, and a failure is never silent.
Google Scholar ──► citations.csv ──┐
profile_stats.json
├──► orig.bib ──► overleaf/Wzmn.bib ──► overleaf/main.tex
papers.csv ────────────────────────┤ (resolved) (+ venue info, (+ \nocite blocks,
(venue, tags, authorship) │ citation counts) citations, h-index)
venues.yaml ───────────────────────┘
identity.json (harvested IDs)
python update.pyEverything else is occasional. update.py exits non-zero on failure and posts a
desktop notification, so an unattended run cannot fail quietly while the CV keeps
looking current.
python update.py --dry-run # show what would change: no writes, no network
python update.py --force # ignore the "nothing changed" checks
python update.py --no-push # build locally, leave the remotes alone
python update.py --skip-fetch # reuse the citation counts already on diskEach step is skipped when the contents of both its inputs and its outputs are
unchanged since it last succeeded (recorded in .pipeline_state.json) — content
hashes, not mtimes, so a fresh clone behaves correctly and a no-op rewrite does
not cascade. Checking the outputs too is what makes a hand-edit heal: revert the
citation totals in main.tex and the next run notices its own output is gone and
rebuilds it, instead of skipping forever because no input changed.
A run where a source never replied does not record its step as done, so the next run asks again rather than freezing a lookup that failed for network reasons.
A failed Scholar fetch does not stop the run. Scholar answers with a CAPTCHA when
it feels crawled, and step 1 refuses to write a short scrape over a good
citations.csv — both are temporary, and neither says anything about the papers
already on disk. So steps 2–7 carry on from what is there, which is exactly what
--skip-fetch asks for on purpose, and the run still notifies and exits non-zero.
The cost is one week of citation counts; the cost of stopping was a paper
published last week staying cited as a preprint because step 3 never ran.
| # | Step | What it does |
|---|---|---|
| 1 | fetch | Scrapes the Scholar profile → citations.csv, profile_stats.json. Time-based (--fetch-age, default 24h). |
| 2 | new papers | Adds anything in Scholar that is not in papers.csv yet. |
| 2b | enrich | Fills blank Authors from Scholar, and resolves a venue from the paper's BibTeX entry. This is what moves a paper out of the CV's ArXiv section once it is published. |
| 3 | resolve | Upgrades arXiv entries in orig.bib to their published version, and looks up rows that have no entry. |
| 4 | build | orig.bib + papers.csv + venues.yaml → overleaf/Wzmn.bib. |
| 5 | tex | Updates \nocite{} blocks, total citations and h-index in overleaf/main.tex. |
| 6 | worklist | Regenerates WORKLIST.md. |
| 7 | push | Commits and pushes to GitHub and Overleaf, rebasing first if a remote moved. |
python scripts/worklist.py # regenerate WORKLIST.md alone (no network)
python scripts/worklist.py --check # exit 1 if anything needs a decision
python scripts/dedupe.py --dry-run # find papers listed twice, keep the published one
python scripts/prune_bib.py # list orig.bib entries nothing refers to (--apply to remove)
python scripts/refresh_venues.py # refresh venue rankings and impact metrics
python -m pytest tests/ -q # the test suitepython scripts/install_schedule.py # weekly local run (macOS launchd)
python scripts/install_schedule.py --show # print the plist / cron line instead
python scripts/install_overleaf_credential.py # store the Overleaf git token, so step 7 can push
python init_new_author.py # wipe personal data, for a forkThe step scripts (fetch_citations.py, build_bib.py, rebuild_tex.py,
resolve_arxiv.py) also run standalone, which is useful when debugging one stage.
git clone https://github.com/borgr/publications.git
cd publications
pip install -r requirements.txt
python init_new_author.py --overleaf-url <your-overleaf-git-url>Note the plain clone: not --recurse-submodules. overleaf/ is a
submodule pointing at the original author's Overleaf project, which you have no
credentials for, so recursing aborts the whole clone with could not read Password. init_new_author.py repoints it at yours and clears the personal
data; run it with --overleaf-url and it does both in one step. Until then the
overleaf/ directory is simply empty, and every step except the push works.
Then edit config.py — that is the only file that is author-specific:
AUTHOR_NAME = "Your Name"
SCHOLAR_USER_ID = "..." # from your Scholar profile URL
S2_API_KEY = "" # optional, see below
CONTACT_EMAIL = "" # optional, for OpenAlex's polite poolcurl and git must be on your PATH.
Editing AUTHOR_NAME is enough on its own: config.py is one of step 5's inputs,
so the next run rewrites the name line in main.tex and repoints the bibliography
styles' name-bolding at you, even if nothing else about the data changed.
Semantic Scholar's unauthenticated access is "1000 requests per second shared among all unauthenticated users" — a global pool, so a long run gets throttled almost immediately. A free key gets "1 RPS" reserved, which is slower on paper but actually completes. It matters more than it looks: the ACL Anthology and OpenReview are both reached through Semantic Scholar, so losing it loses three sources.
Request one at https://www.semanticscholar.org/product/api, then put it in a file outside the repository:
mkdir -p ~/.config/publications
printf '%s' 'YOUR_KEY' > ~/.config/publications/s2_api_key
chmod 600 ~/.config/publications/s2_api_keyNot in config.py, even though it has a slot for it: that file is tracked and
this repository is public, so a key there is one git add -A from being
published. tests/test_no_secrets.py fails the build if one ever is. The slot is
kept for a fork that keeps its config private.
The file, rather than an exported S2_API_KEY, because launchd hands a scheduled
job PATH and HOME and none of the shell's exports — a key in .zshrc would
cover every run except the weekly one that does the most lookups. Order of
precedence is environment, then this file, then config.py; the environment comes
first so CI can inject a key for one run without writing it to disk. Each run
prints which of the three it is using, since a key being shadowed by another
source is otherwise invisible.
Without a key everything still works, with more waiting: requests to Semantic Scholar are spaced 3s apart instead of 1.2s, and a refusal pauses it for two minutes instead of one.
Worth knowing if you fork this and point it at a different set of papers, because the failure it prevents does not look like a failure.
Every request is spaced and paused per host. A refusal from one source used to put every source behind the same cooldown, so a busy Semantic Scholar silenced OpenReview too — and since a request that was never made is recorded as "no answer", not as "no published version", the run would honestly report that it had learned nothing about a paper whose ACL entry was sitting right there.
Semantic Scholar is asked about every arXiv ID in the run in one batch request, before any paper is resolved. This is not only a courtesy. On a real weekly run — with a key, and 1.2s between requests — Semantic Scholar refused twice, went into cooldown, and 32 of the remaining 34 lookups were never attempted: it contributed nothing to that run, and took the ACL Anthology and OpenReview down with it. Spacing requests further apart does not fix a budget; making fewer of them does. The batch runs with or without a key, since one request is easier on the shared anonymous pool than thirty-four.
A source that keeps serving a refusal — DBLP answers rate limiting with an HTML error page and HTTP 200 — is paused for the rest of the minute rather than asked again, with backoff, for every remaining paper. A request that never completed is treated as the network rather than the source, and pauses nothing.
A pause is then waited out rather than skipped, up to 150s per host per run. The pause is measured in wall-clock time and the loop it protects finishes in seconds, so the two do not overlap the way they look like they should: on a real run Semantic Scholar's search endpoint went into a 60s pause with nine papers left, all nine were skipped inside two seconds, all nine were reported unknown — and the next run does exactly the same thing, so those papers never resolve at all. Waiting is also the more polite of the two options: nothing is sent while the source has asked not to be asked, and then the same requests go out that would have gone out anyway. The budget is what stops a source that refuses everything from costing the run an afternoon; once it is spent, that host is skipped again.
Each refusal also doubles that host's spacing for the rest of the run, capped at 10s, because resuming at the pace that was just refused spends the waiting budget on being refused again. A 429 is the source saying the current rate is too fast for it now, for reasons no constant in this repository can know — the run that prompted this was refused twice at 1.2s spacing while holding a key documented at 1 RPS.
Split deliberately, because Scholar blocks datacenter IP ranges — a hosted runner gets a CAPTCHA, not data:
- Locally, weekly (Monday 08:37), via
scripts/install_schedule.py. This is the one that fetches from Scholar. It also needs the Overleaf token stored once, withscripts/install_overleaf_credential.py— see below. - GitHub Actions covers what does not need Scholar: the test suite on three Python versions, the oldest supported dependency versions, a rebuild from committed data, a determinism check, a fork-from-scratch check, and a failure if the table has duplicate rows or an ambiguous citation join. No secrets, works in a fork. It also runs weekly, on Monday at 14:47 UTC — deliberately after the local run, since checking first would compare against data an hour from being replaced.
Step 7 pushes with plain git push, so git has to be able to authenticate on its
own — there is nobody at the keyboard at 08:37 on a Monday. Run this once, and
again whenever you rotate the token:
python scripts/install_overleaf_credential.pyIt takes the URL from the first of three places that has one:
~/.config/publications/overleaf_git_url, a one-line hand-over file. Use this when the person holding the token and the person running the installer are not the same, or not at the same keyboard. It is read, transferred into the credential store, verified, and then deleted; a token it could not authenticate with is left in place so you can correct it. It lives outside the repository deliberately, and the installer refuses to read a path inside the repository rather than riskgit add -Acommitting a token to a public remote.$OVERLEAF_GIT_URL, if you have already exported it.- A hidden prompt.
From there it goes to git's own credential store — the macOS keychain, via
git credential approve — and a real request proves it works. The token is never
written into the working tree, never put on a command line, and never printed;
storing it in the submodule's remote URL or in .git/config would leave it in
plaintext, and a dotfile in the tree is one git add -A away from a public repo.
The store is then read back with git credential fill, because approve
cannot be trusted to have worked: git exits 0 whether the helper stored anything
or not, so a credential.helper naming a program that is not installed
(libsecret where git-credential-libsecret was never built) reports success
having stored nothing. The read-back also catches a different token answering
first — git uses the first helper that answers, so an old token ahead of the new
one is the one git push sends.
Until you do this, update.py prints the reason it will not be able to push
before step 1 rather than after a full Scholar fetch, and then runs anyway.
Only the push is lost: papers.csv, citations.csv and the rebuilt CV all end up
current on disk and committed locally, and step 7 fails, notifies and exits
non-zero on its own. Tokens get revoked and rotated, and one that froze the whole
pipeline would quietly stop the data tracking reality as well — a worse failure
than a stale Overleaf. GIT_TERMINAL_PROMPT=0 keeps a missing credential from
stalling on a password prompt an unattended run has no way to answer.
Without this, nothing outside your own machine can tell whether the CV Overleaf
compiles matches the committed data — the failure that motivated it was a run of
--no-push runs that left Overleaf hundreds of citations behind while every run
reported success.
-
In Overleaf: Account Settings → Git integration for a token, and Menu → Git for the project URL. They are separate — the URL Overleaf shows you carries no credential, and cloning it in CI fails asking for a password.
-
Combine them into one URL with the token as the password, and store that as a GitHub repository secret named
OVERLEAF_GIT_URL(Settings → Secrets and variables → Actions → New repository secret):https://git:YOUR_TOKEN@git.overleaf.com/YOUR_PROJECT_IDUse a token generated for this, not whatever your local clone authenticates with, so revoking CI's access does not break your own pushes.
CI then rebuilds the CV from the committed data and fails if Overleaf would
compile something different. To have CI push the fix rather than just report
it, add a repository variable PUBLISH_TO_OVERLEAF set to true. That is off
by default: Overleaf is a document you also edit by hand, so writing to it is a
decision, not a default. It rebases before retrying, so a push and a hand edit
racing does not fail the run.
Publishing runs on the default branch only, and one run at a time. Overleaf holds
one CV, so a work-in-progress branch has nowhere good to put its version of it:
publishing would overwrite the real one, and failing would go red for exactly the
work the branch exists to do. Two runs pushing at once is the other half of that —
the job takes a concurrency lock rather than racing.
With no secret set, the job prints how to enable itself and passes — a fork is
never red for a project it does not have. The same staleness is reported locally
in WORKLIST.md, which needs no credentials.
The token is never in the repository, so forking this project cannot hand anybody your Overleaf account. Concretely:
- Nothing committed carries a credential.
tests/test_no_secrets.pyfails the build if one ever is, so this is checked rather than asserted. - Secrets do not travel with a fork. They live in your repository's settings, encrypted, and GitHub will not read one back to you either — only overwrite it. A fork starts with none, which is why the job is written to pass when it finds none.
- A pull request from a fork gets no secrets, by GitHub's design, so a PR
that edits the workflow to print the token cannot: the job is also skipped on
pull_requestoutright. - Logs are scrubbed. GitHub masks any registered secret's value in output,
the URL is passed only through the environment, and the clone step redacts
//…@out of git's own error text before printing it.
What is public is the project URL in .gitmodules
(https://git@git.overleaf.com/<project-id>). That is an address, not a key:
cloning it without the token fails, which is exactly what the plain-clone
instruction above is about.
| File | Purpose |
|---|---|
papers.csv |
Source of truth. One row per paper: venue, authors, year, BibTeX key, and the tag flags that drive CV sections. Opens in Excel/Numbers/LibreOffice. |
venues.yaml |
Venue rankings, impact metrics, and the sentence each venue prints. |
config.py |
Author name, Scholar ID, optional API keys. |
orig.bib |
Curated BibTeX. Mostly maintained by step 3; hand-edit what it cannot resolve. |
Generated but committed, because each is expensive to rebuild or useful to
browse: citations.csv, profile_stats.json, identity.json (harvested
identifiers), resolve_attempts.json (retry counters), .pipeline_state.json
(the input and output hashes each step last saw), WORKLIST.md,
overleaf/Wzmn.bib.
papers.csv is the only table format. table_io.py addresses every column by
header name — so inserting or reordering one cannot misfile a value — and
validates the whole table on every load, because a spreadsheet round-trip that
reformats a column produces a table that still builds, just wrongly.
| File | Purpose |
|---|---|
update.py |
The orchestrator. |
fetch_citations.py |
Scholar scraper, including each paper's stable citation_for_view ID. |
table_io.py |
Reads/validates/writes papers.csv, by column name. |
identity.py |
Stable identifiers, and the joins built on them. |
citations_io.py |
citations.csv reading and writing. |
resolve_arxiv.py |
arXiv → published BibTeX, via DBLP / S2 / ACL / OpenReview / DOI / OpenAlex. Finds candidates; writes nothing. |
bib_edit.py |
Decides what may change in orig.bib and writes it: keys, the venue transplant, the in-place update. |
build_bib.py |
Builds Wzmn.bib and assigns each paper to a CV section. |
rebuild_tex.py |
Updates main.tex in place. |
bib_utils.py |
Brace-counting BibTeX parser, text normalization, publication ranking. Reads only. |
venues.py |
Loads venues.yaml. |
pipeline_state.py |
Content-hash step skipping, over inputs and outputs both. |
notify.py |
Failure notification (macOS Notification Center, Actions annotation). |
scripts/papers_fig.py, scripts/papers_graph.py |
Standalone figures. Not part of the pipeline; needs requirements-figures.txt. |
Titles are not stable: BibTeX braces capitalization ({B}aby{LM}), Scholar
lowercases subtitles, papers get retitled between preprint and publication, and
the table is typed by hand. So matching is by identifier, and titles are only
the bootstrap:
- Stable ID. Scholar's
citation_for_viewper paper;externalIdsfrom Semantic Scholar for ArXiv/DOI/ACL/DBLP together — the crosswalk for when the ACL record knows no arXiv ID and vice versa. Recorded inidentity.jsonthe first time it is seen, so a fuzzy match happens at most once per paper. - Normalized title. Case, punctuation and BibTeX braces stripped. An identity claim, not a guess.
- Fuzzy title, only with a clear margin over the runner-up, and reported.
Anything ambiguous is reported, never guessed. Step 2's "is this paper already known?" calls the same matcher as the citation join, deliberately: when they disagreed, step 2 appended a duplicate row for every fuzzy-matched paper on every run.
Two Scholar records for one paper are summed — Scholar splits a paper across records until you merge the versions, and each record counts different citing papers. Records that are not the same paper are never summed.
bib_utils.publication_rank and choose_published are the single rule, used by
step 3 (never downgrade), by the build (emit the version of record), and by
scripts/dedupe.py:
- published beats preprint — the version of record is what a CV should cite;
- within the same class, the newer year wins, because two preprints of one paper are its v1 and v2 and the newer carries the current title;
- then publication rank, content length and key, so the result is stable.
Duplicates are found by identifier as well as by title, which catches the retitled ones no title comparison can.
A duplicate group is a set of table rows — (title, key) pairs from
table_io.rows_named — not a set of titles. The commonest duplicate is one title
entered twice, and a group of titles collapses that to a single member: the build
then ranked the row's entry against itself, declared it the winner and suppressed
it, so the CV printed the arXiv version of a paper whose ACL entry was sitting in
the table, with a note reading "published beats published". dedupe reported the
same pair as nothing to fix, and its removal matched rows by title — which for two
rows sharing one title deleted both, the merge removing the paper it was merging.
A mis-resolution is a different problem from a duplicate, and is never fixed
automatically: a duplicate is provable and safe to drop, while an entry pointing at
the wrong paper has to be looked at. The cheap test is whether the entry credits
the author at all — bib_utils.lists_author — which is what step 3 applies to
every candidate and what tests/test_author_on_every_paper.py asserts over the
whole table. Without it the resolver accepted an invited talk by one of a Nature
paper's twenty co-authors as that paper's published version, on a title similarity
of 0.86, and the CV printed it.
Every source in the ladder answers a question it was asked, and three of them will
answer it with something else if they have nothing: DBLP, Semantic Scholar and
OpenAlex all run relevance searches, ranked, always non-empty. So an entry is only
accepted when it is a version of the paper that was asked for —
resolve_arxiv.titles_agree: the two titles have to negate the same things,
normalized similarity has to be at least 0.72, and short of 0.95 the year has to
agree too, since difflib rewards a long common subsequence however much extra text
there is.
The negation test is there because character similarity has no notion of meaning and the four characters that reverse a claim cost nothing: "Attention is all you need" and "Attention is not all you need" are 93% identical, a year apart, and a paper and a rebuttal of it. Titles like that are common enough here that a search for either can return the other.
Resolving by identifier is exact, which is not the same as being right — the
identifier can be wrong, and then what comes back is a real, well-formed,
verifiable entry for somebody else's paper. So the DOI, ACL Anthology and
OpenReview rungs check their entry as well, not just their lookup. Every one of
those identifiers arrives from somewhere: externalIds on an S2 record, a store
record an earlier run wrote, a URL in the existing entry.
One rule in one function, applied at every point where an entry is produced. OpenAlex used to keep a private second copy of the idea — a bare 0.90 similarity, with no year and no negation test — which made it both stricter and blinder than the shared rule: it would take a 0.92-similar paper from another decade, and reject the retitled journal version it is in the ladder to find.
The single rule matters because it took only one unguarded rung to publish the wrong paper: an unchecked S2 title search answered "Every eval ever: Toward a common language for AI eval reporting" with a Lancet epidemiology paper, its DOI became that paper's known DOI, clibib resolved it exactly, and the CV listed "Global burden of 292 causes of death in 204 countries and territories". Nothing raised. The two titles score 0.22; a genuinely retitled paper scores 0.77 and passes on its year.
orig.bib accumulates: a paper's arXiv entry stays behind when step 3 moves its row
to the published key. Sixty-nine of a hundred and seventy-eight entries were
unreachable that way. None of it reaches the CV — Wzmn.bib is built from the
intersection of the table and the bibliography — so this is about being able to read
the file and see a real diff in it.
scripts/prune_bib.py reports them, and removes them with --apply. Three things
protect an entry, and any one is enough: a table row's Bib key, a \nocite in
main.tex (including a commented-out one), or a Scholar ID bound to it in
identity.json. It also strips pretitle fields committed back into the source,
which every build regenerates from papers.csv anyway. Git history is the backup.
venues.yaml holds each venue's kind (journal, conference, or other for a
real outlet with no ranking, like a blog), its description (the sentence the CV
prints), and match: phrases used to recognise it in a raw venue string.
scripts/refresh_venues.py refreshes the numbers from Google Scholar Metrics and
OpenAlex. A venue marked manual: true keeps its prose; only its metrics block
is updated. When this file was created, four of eleven rankings were wrong — EMNLP
and NAACL both claimed to be 1st.
A venue the pipeline cannot place gets no venueinf line and its paper is filed
under ArXiv Articles, so unplaceable venues are reported in WORKLIST.md. Add a
match: phrase to fix one.
clibib resolves an identifier to BibTeX via a
Zotero translation server. pip install clibib and step 3 will use it for
DOIs, covering journals and book chapters that DBLP and the ACL Anthology do
not index. It resolved 10 entries on a real run. Without it the pipeline behaves
exactly as before.
Deliberately never used for title search: measured against this repo's own unresolved papers, its free-text lookup returned a confidently wrong paper 2 times in 5 with no error raised. Its identifier paths are exact and fast, and its own README recommends preferring them.
What comes back is still checked against the title that was asked for, because an exact lookup of a wrong DOI returns a wrong paper just as confidently — see An answer is not a match above.
As a manual helper it is genuinely useful for worklist items where you have an identifier:
clibib 10.1038/s41586-021-03215-wEach entry in Wzmn.bib carries a citations={N} field, which the BST emits as
\bibcitecount{N}. Toggle display in main.tex:
%\newcommand{\bibcitecount}[1]{ \textit{\small[#1 cited]}} % show
\newcommand{\bibcitecount}[1]{} % hide (default)overleaf/ is a git submodule pointing at the Overleaf project. Step 7 pushes
main.tex and Wzmn.bib there, then pushes this repo to GitHub. If the Overleaf
remote has moved — because you edited the project in Overleaf's own editor — the
push rebases onto it and retries, so that no longer needs a manual pull.
Pull Overleaf-side edits back with git -C overleaf pull origin.