Skip to content

v0.5.0

Choose a tag to compare

@goerz goerz released this 24 Jul 03:39
· 49 commits to master since this release
v0.5.0
  • Changed: the import command now reads its source file from a --file FILE option instead of a positional FILE argument, so a positional argument ending in .bib always names the library, like every other command. This removes the collision where, with a default_bib_file configured, bibdeskparser import from_paper.bib silently claimed the snippet as the library and then failed for lack of a source. Migration: bibdeskparser import library.bib entries.bib becomes bibdeskparser import library.bib --file entries.bib, and bibdeskparser import from_paper.bib (into the default library) becomes bibdeskparser import --file from_paper.bib. Library.import_bibtex, which takes the BibTeX text directly, is unaffected. [[#46], [#51]]
  • Added: a read-only config CLI command that dumps the resolved configuration -- the built-in defaults merged with whatever a discovered bibdeskparser.toml sets. Unlike config_path (which reports only the file in effect, and fails when none is found), config shows the effective value of every setting, including the ones a file omits (auto_key.clean, preprint_export, the built-in preprint_archives, ...) and the built-in defaults in full when no file exists. The default text output is TOML-shaped, mirroring what a bibdeskparser.toml would contain to reproduce the resolved tunable state (an unset value, or an auto_key/auto_file table without a format_spec, is omitted); --no-types restricts the dump to those user-tunable settings, while the default additionally lists the resolved entry-type/field data model (documented_types, recognized_entry_types, universal_fields, known_fields), and --json prints the complete state as an object with unset values as null. The BIBFILE argument is optional -- it only fixes the config-discovery directory, defaulting to the current directory -- so the command needs no .bib file and never fails for a missing configuration file. [[#45], [#50]]
  • Added: a --key-format option on the check CLI command, an opt-in audit (off by default) that reports every citation key not matching its expected auto-key format, i.e. every key that eval_format_spec (or single-argument rekey) would regenerate differently. A preprint-only entry is audited against the arXiv preprint format, every other entry against the configured auto-key format; a --format-spec PATTERN option audits against that pattern instead and implies --key-format (combining it with --no-key-format is an error). A key already matching the format evaluates to itself, so disambiguated sibling keys such as SmithPRA2015 and SmithPRA2015a both pass; an entry that lacks a field the format requires is reported as unevaluable rather than silently skipped; and when no format is available at all (no --format-spec and nothing configured), a single message is reported instead of one failure per entry. The audit name key_format also appears in the --json output. [[#44], [#49]]
  • Added: a --files option on the check CLI command, an opt-in audit (off by default) that reports every linked attachment (bdsk-file path) that does not resolve to a real path on disk relative to the .bib directory. It walks each stored path one component at a time, matching case exactly, so besides a link to a deleted file it also catches a link whose spelling differs only in case from the file on disk, one that works on a case-insensitive filesystem (macOS) but breaks on a case-sensitive one (a collaborator's machine or a Linux CI job); the plain existence check behind the warning a write-in-place command prints for a missing link cannot detect that case-mismatch class, so check --files can fail a library that would be written without any such warning. It is off by default because attachments may legitimately live only on another machine, so a fresh clone of a library whose PDFs are not under version control would otherwise fail wholesale. The audit name files also appears in the --json output; an attachment whose stored path is empty is flagged, and a link resolving to a directory passes (BibDesk can link folders). [[#43], [#48]]
  • Added: a --usage option on the bibdeskparser command-line tool, printing a short usage summary (the one-line description, the usage line, and the list of command names) as a compact alternative to the full --help output. [[#47]]
  • Changed: running bibdeskparser with no command now prints the short usage summary (see --usage) on stderr and exits 2, instead of dumping the entire --help output. [[#47]]
  • Fixed: the network commands no longer fail in an environment that routes traffic through a SOCKS proxy (ALL_PROXY=socks5h://..., e.g. an SSH tunnel). Both HTTP stacks used by the package refuse to even attempt a SOCKS connection without an optional helper package, each with a different error: httpx (behind add and import --url) needs socksio, and requests (behind add_preprint, add_doi, and add_abstract, via the arxiv and habanero packages) needs pysocks. Both helpers are now regular dependencies.
  • Fixed: with --dry-run, the per-key reports of the add_abstract, add_preprint, and add_doi CLI commands (and the preprint report of add --add-preprint) now say would store / would mark known missing instead of the past-tense stored / marked known missing, which misread as the file having been modified. The JSON reports are unchanged: applied marks what a real run would store.
  • Changed: the command-line tool now reports save-time warnings as the same clean Warning: lines on stderr it uses for all other warnings, instead of letting them surface through Python's warning machinery with a meaningless cli.py source location. Warnings about linked files that do not exist are printed individually only up to five; beyond that, they collapse into a single summary line with the total count and the first missing file. Previously, saving a .bib file separated from its attachment tree (e.g. a copy in another directory) flooded stderr with one location-prefixed UserWarning per linked file, on every mutating command.
  • Fixed: importing an entry whose journal spells out a name the library abbreviates no longer aborts with a macro-name collision (publisher and Google Scholar exports spell journals in full, e.g. Physical Review Letters, or Scholar's lowercased Physical review letters, where the library defines prl = "Phys. Rev. Lett."). When the derived macro name is taken, the incoming name is compared word by word against the colliding macro's value -- a dot-terminated word matches as a case-insensitive prefix, a bare word must match exactly, so Phys. Rev. A does not capture Physical Review Applied -- and on a full match the existing macro is reused, with a warning showing the journal_macros alias line that makes the mapping explicit. When the match fails (the colliding macro holds a different journal), the error now prints the exact journal_macros configuration line for each possible intent -- appending the incoming spelling to the alias list (canonical value first), or an entry under a fresh macro name / an initials.journal exception -- as does the error for a journal_macros entry that conflicts with an existing @string definition. [[#40], [#42]]
  • Added: a with_files argument of Library.keys and a paired --with-files/--without-files option on the keys CLI command, a tri-state filter on file attachments (the bdsk-file-N fields): the default (None, or neither flag) does not filter, True/--with-files keeps only entries with at least one attachment, and False/--without-files only those with none. This composes with the existing keys filters, so bibdeskparser keys --type article --without-files lists the articles still needing a PDF. [[#39], [#41]]
  • Changed: the files, urls, groups, and keywords CLI commands now share one output shape. Each takes any number of citation keys and prints a map from each citation key to its list of values (attachment paths, URLs, group names, or keywords): with keys, exactly those entries (an entry with none maps to an empty list); with no key, every entry in the library that has at least one value, in library order. A new paired --flat/--no-flat option (default --no-flat) instead prints just the values as a bare list, combined across the selected entries with duplicates removed (its order is unspecified); files --flat is thus every file the library references, the reverse index for reconciling against a folder of PDFs. files keeps its --absolute/--relative option; groups and keywords gain a paired --index/--no-index option that prints the inverse map instead, from each static group or keyword to the citation keys it contains (--index takes no keys and does not combine with --flat). Migration: a single-key files KEY / urls KEY / groups KEY / keywords KEY that previously printed a bare list now prints a one-entry KEY: values map; pass --flat for the old bare-list output (single-key files --flat keeps bdsk-file-N numeric order). A bare groups / keywords (no key) previously printed the group/keyword catalog and now prints the per-entry map; pass --index for the catalog. [[#39], [#41]]
  • Fixed: Library.add_preprint (and the add_preprint CLI command) no longer reports a false no-results for an entry whose title or first-author last name contains a Latin letter that Unicode treats as atomic rather than as base-plus-combining-accent (ø, Ø, ł, Ł, ß, æ, œ, đ, ð, þ, ħ, ı, ŋ, and friends) -- the canonical case being any title naming the Mølmer-Sørensen gate. Such letters have no Unicode decomposition, so the old accent-folding left them in place and the ASCII-only tokenizer then split the word on them (sørensen became s/rensen), poisoning every arXiv query and silently disabling the title+author acceptance rung for affected first authors. arXiv indexes title characters literally, so queries now preserve these letters intact, while match comparison folds them to ASCII (so an ASCII-spelled publisher title still matches the unicode arXiv record, and vice versa); letter-producing TeX commands in a raw title (e.g. {\o}) are also decoded before a query is built. With a known-missing group configured for eprint, the false negative had additionally marked the entry as verified-absent, excluding it from future searches. [[#36], [#38]]
  • Changed: the names audit of the check command now also flags an author/editor whose parsed first name has a part that cannot be initialized -- a hyphen-separated segment that does not begin with a letter after TeX-to-unicode conversion, such as a quoted nickname (`Eunice') copied into the author list, or a stray hyphen that detaches an initial (Meyer, H -D or Meyer, H- D, both of which should render H.-D.). Such values split cleanly into names, so they passed the old audit, yet they make render emit a bogus initial (Y. K. `. Lee) or silently drop the hyphen (H. D. Meyer); the gate now reports them for manual repair rather than letting the corruption surface only in rendered output. [[#33], [#37]]
  • Added: Library.add_doi and a corresponding add_doi CLI command, recording the DOI of an existing entry in its doi field -- either an explicitly given DOI (--doi, validated and normalized to its bare lowercase form, no network access), or one found online: the DOI recorded on arXiv for the entry's eprint (which names the published version of exactly this paper; an arXiv-issued 10.48550/... DataCite DOI does not count), or otherwise a Crossref search by title and first author. A search result is stored only on a confident match (a near-exact title, or a good title corroborated by the first author's last name); a title-based match whose publication year differs from the entry's year by more than one is rejected as a likely title collision (reported as year-mismatch for manual review), and an erratum, corrigendum, retraction, comment, or reply never matches an entry that is not itself such an amendment. Preprint-only entries are skipped (the search would find the published version's DOI, which does not belong on a preprint reference). With a known-missing group configured for doi in the known_missing table of bibdeskparser.toml (e.g. doi = "No DOI"), the group is maintained exactly as add_abstract/add_preprint maintain theirs: members are skipped, a clean no-match marks, storing a DOI unmarks, and overwrite/--overwrite re-audits; entries verified to have no DOI thus also pass the missing-doi audit of check automatically. [[#35]]
  • Added: --group NAME and --not-group NAME filter options on the keys CLI command, and corresponding group/not_group arguments of Library.keys, keeping only entries that are (respectively, are not) members of the given static groups (repeatable; group names are matched case-sensitively, and an unknown group name is an error rather than an empty result, so a typo cannot silently select nothing, or everything for --not-group). For example, re-audit the entries recorded as having no preprint with bibdeskparser add_preprint --overwrite $(bibdeskparser keys --group "No Eprint"). [[#34]]
  • Added: a known_missing table in bibdeskparser.toml, mapping a field name to the name of a BibDesk static group that records, per entry, a verified "searched, this info does not exist" status (e.g. abstract = "No Abstract", eprint = "No Eprint", doi = "No DOI"; exposed as Library.config.known_missing). This replaces the previous empty-field markers, which the BibDesk app silently deletes whenever it saves the .bib file; static groups survive BibDesk saves, BibDesk maintains their membership when citation keys change or entries are deleted, they never appear in exported entries, and they can be managed by drag and drop in BibDesk. With the table configured, Library.add_abstract/Library.add_preprint (and the corresponding CLI commands) skip entries in the field's group (reported as known-missing; overwrite/--overwrite re-searches), add an entry to the group when a search runs cleanly and finds nothing (creating the group on first use; a failed search never marks), and remove the entry from the group whenever a real value is stored; without the table, none of this bookkeeping happens. Membership in the group configured for doi makes the check command accept an article without a doi. [[#34]]
  • Added: two new check audits: empty_fields flags every defined-but-empty field on any entry (BibDesk deletes empty fields when saving, so the field would silently disappear the next time the library is saved in BibDesk), and known_missing flags an entry that is a member of a configured known-missing group while actually having a non-empty value in that field (e.g. after a manual edit in BibDesk). [[#34]]
  • Changed: a defined-but-empty field now counts as missing everywhere, and the empty-field "audited" markers are gone. keys --empty and the empty argument of Library.keys are removed (--missing/missing now also match a defined-but-empty field); the mark_empty arguments and --mark-empty options of add_abstract/add_preprint, the mark_empty key of the add_abstract table, and the add_preprint table (whose only key was mark_empty) are removed; a defined-but-empty doi no longer suppresses the missing-doi check problem (membership in the known-missing group configured for doi does instead); set_field KEY FIELD "" is now an error (use delete_field, or record a verified absence via a known-missing group); and applied in the add_abstract/add_preprint results now means the library was modified (the field, or the known-missing group membership). To migrate a library that used the old markers, run bibdeskparser check: every leftover empty field is reported by the new empty_fields audit. For each reported entry, record the verified absence in a group (bibdeskparser set_group "No Eprint" once to create the group, then bibdeskparser add_to_group "No Eprint" KEY...), add the matching known_missing entry to bibdeskparser.toml, and remove the empty field (bibdeskparser delete_field KEY eprint; likewise for abstract and doi). Also delete any mark_empty key and add_preprint table from bibdeskparser.toml. [[#34]]
  • Removed: Entry.add_abstract and Entry.add_preprint. Use Library.add_abstract(key, ...) and Library.add_preprint(key, ...) instead: the known-missing group bookkeeping and the PDF-attachment lookup both need the library, so the Library methods are the only public fetching entry points. Entry.add_abstract's pdf_path argument is gone with it (the Library method locates the entry's first attached PDF itself). [[#34]]
  • Changed: when no abstract is found and one of the consulted sources failed (an unreachable API, or an attached PDF that could not be read), Library.add_abstract now returns source="error" instead of "none", mirroring add_preprint's match="error"; the negative is not conclusive, so it never records a verified absence. Previously a network failure was indistinguishable from a definite "no abstract exists anywhere". [[#34]]