Skip to content

Releases: tamnd/papers-reader

v0.8.25

Choose a tag to compare

@github-actions github-actions released this 19 Sep 11:34
v0.8.25
d596a91

What's Changed

  • audit: report a paragraph cut in the middle of a word by @tamnd in #237
  • translate: repair a lost dollar pair and a translated name by @tamnd in #238

Full Changelog: v0.8.24...v0.8.25

v0.8.24

Choose a tag to compare

@github-actions github-actions released this 19 Sep 11:04
v0.8.24
3db3c93

What's Changed

  • translate: accept a repeated formula and put back bare dollars by @tamnd in #236

Full Changelog: v0.8.23...v0.8.24

v0.8.23

Choose a tag to compare

@tamnd tamnd released this 19 Sep 10:22
v0.8.23
d4f9328

A translate run stops itself after three files in a row fail, on the reading that three failures that close together mean the fleet is down. With several jobs in flight over theory papers that reading is wrong, and the Vietnamese pass over the corpus stopped after 86 of 862 files because Razborov, Goldwasser and Shamir each lost a chunk to the mathematics check at the same time.

A chunk that is never accepted now comes back as a typed refusal, and only a failure that is not a refusal counts towards stopping the run.

What's Changed

  • audit: T08 reads a short section against the paper around it by @tamnd in #233
  • translate: do not read a refused page as a fleet outage by @tamnd in #235

Full Changelog: v0.8.22...v0.8.23

v0.8.22

Choose a tag to compare

@tamnd tamnd released this 19 Sep 10:03
v0.8.22
8cbdc15

One fix.

A figure's region grows up from its caption until it meets the paper's own prose, and which lines counted as prose was decided by asking whether anything stood beside a line anywhere on the page. On a two column page nothing above the figure is ever alone, so the region ran past all of it and the prose went into the PNG. The question is now asked of the column the line is set in, counting only what stands above the caption, and the whole page reading is kept as the fallback for a figure drawn out of type, where the column by column reading collapses the region.

Over the corpus this takes the manifest from 502 figures to 505, loses none, and tightens 42 of them from a median 336 points tall to 134. Nothing grew.

What's Changed

  • figures: judge the prose above a figure column by column by @tamnd in #232

Full Changelog: v0.8.21...v0.8.22

v0.8.21

Choose a tag to compare

@tamnd tamnd released this 19 Sep 08:04
v0.8.21
b466fe8

One new audit rule.

L22 reads a translated file looking for a number with an English scale word after it, a thousand or a million or a billion left in place where the language has its own word. It is a hard rule, because a reader in Vietnamese who meets "1 million" in the middle of a Vietnamese sentence is reading English.

Over the whole corpus it named one file, the acknowledgements of the Gamma database machine paper, which has since been asked for again and now reads "1 trieu".

What's Changed

  • audit: L22 finds an English scale word left after a number by @tamnd in #206

Full Changelog: v0.8.20...v0.8.21

v0.8.20

Choose a tag to compare

@tamnd tamnd released this 19 Sep 06:38
v0.8.20
a6609ae

A section file with an empty body publishes nothing. It carries a title in its front matter, it takes a number in the run of files, and a reader who opens it is shown a blank page. Three papers were producing nine of them and none of the nine was ever committed.

Every one came from a heading the reader set over another heading with no prose in between. Saltzer's section V is a heading and then its first lettered subsection. The Turing award lecture on clocks heads its appendix and names the proof in it on the next line. Sketchpad's appendix G says the distinctive features of TX-2 are, and the reader set every item of the list that follows as a heading of its own.

The splitter now gives such a heading to the section under it, as a subheading, because a heading is a heading of what comes after it. The file numbers are given out after the fold, so a section that has gone does not leave rule T04's hole in the numbering. A repeat of the paper's own title with nothing under it is a running head and goes nowhere.

Rule T14 is new and hard: no section file has an empty body. T08 reported these already and reported them softly, along with the one line sections and the splits that landed two words early.

The prune report is fixed here too. It said every stale file had been deleted whenever --prune was given, and a partial split deletes only the files that collide with a section number this split just wrote, so the note claimed deletions that had not happened.

What's Changed

  • A heading with nothing under it is not a section by @tamnd in #231

Full Changelog: v0.8.19...v0.8.20

v0.8.19

Choose a tag to compare

@tamnd tamnd released this 19 Sep 06:21
v0.8.19
0312f2c

The hard audit over the corpus now finishes with nothing to report: 76 rules, 71 passed, none failed and 5 not run. This release is the last three rules that were still failing, on top of the six repairs that came before them.

R08 was the Karp reprint. The PDF carries the two books the editors of the collection suggest reading under a heading that says References, and under that heading the reprinted article starts, so the index came back with 19 of Karp's own numbered combinatorial problems read as references. The bibliography now ends where the splitter says the next section begins, which is the same question R08 asks. It also takes the appendix off the end of the last entry in MapReduce, ResNet and Saltzer.

M11 was a display that never closes. A sentence broken across a page arrives with no capital at the front of it, a reader looking at that opens a display, and the display has nothing on the page to close it. extract.Undisplay writes those back as prose, with the punctuation moved out of the formula. An unclosed display is the whole of the test, so a real display is never touched, and an unclosed display with no words in it is left for M11 to report.

T04 was a page that had been refused three times and then lost. The GNMT mixed word and character model prints its markers in angle brackets, b is the bold tag, and acceptance rule A10 read the page as raw HTML. Not one of the tags in the two thousand extracted pages is written in capitals, so a token of one letter in capitals is a word and not markup. The predicate is exported now and A10 and audit rule T11 both call it, because they are the same question at two moments.

What's Changed

  • Take the last of the mathematics out of the listings by @tamnd in #229
  • Audit repairs: mathematics rules, a partial prune, and whole bibliographies by @tamnd in #230

Full Changelog: v0.8.18...v0.8.19

v0.8.18

Choose a tag to compare

@tamnd tamnd released this 19 Sep 00:08
v0.8.18
099f15c

Three more kinds of heading the splitter used to walk past.

A roman numeral does not need a full stop after it. The RSA paper heads
its sections "I Introduction" and "V Our Encryption and Decryption
Methods", so nothing read a number off any of them and the paper came
out as three files cut at two display formulas. The stop is now
optional, and a stopless numeral has to be followed by a title that
starts with a capital before it counts, which keeps a line of prose
that opens with the word I out of it.

A display formula is not a heading. A formula on a line of its own is
short, all capitals and has no terminal punctuation, which is what the
typographic detector looks for, and two of them became section titles
in the same paper. A paragraph that opens with a display delimiter is
refused now. Mathematics inside a real heading still works.

A heading found by name can be run into the paragraph under it.
Fourteen papers printed References on the same line as the first entry
of the bibliography, which filed the whole reference list inside the
conclusion. The name table is stronger evidence than the leading digit
the run-on pass used to insist on, because a line that reads exactly
"References" is not a line of anybody's prose.

The last one is the bibliography that prints its heading again at the
top of each of its pages. The splitter cut at every repeat, so the
congestion avoidance paper had its twenty five entries spread over
three files and a reader following a citation had to guess which. The
repeats are taken out when there is nothing but the entry list between
them, so a paper that really does print two bibliographies still gets
two sections.

RSA went from 3 sections to 12. The corpus went from 34 papers with no
references section to 20.

What's Changed

  • Find three more kinds of heading the splitter was missing by @tamnd in #228

Full Changelog: v0.8.17...v0.8.18

v0.8.17

Choose a tag to compare

@tamnd tamnd released this 18 Sep 23:42
v0.8.17
21557d1

Two papers came out of the split with no front matter, and this gives it
back to them.

The RSA paper has no numbered sections that the splitter recognises, so
the appendix scan had no last heading to start after and began at
paragraph one. It read the paper's own title as appendix A, which took
the title away from the front matter. The scan now returns nothing when
there are no body headings at all, because an appendix comes after the
body and a paper with no body has nowhere to put one.

The AlphaGo paper prints the word ARTICLE over the title, and a reader
transcribes that as a heading like everything else set large. The title
was only compared against the first heading, so the masthead stayed and
took the title, the authors and the abstract into section one. The title
is now looked for over the first three headings, and the search stops at
the first heading that carries a number of its own so that a paper whose
section one is named after the paper keeps it.

What's Changed

  • Give two papers back their front matter by @tamnd in #227

Full Changelog: v0.8.16...v0.8.17

v0.8.16

Choose a tag to compare

@tamnd tamnd released this 18 Sep 23:26
v0.8.16
4aae0c9

Two acceptance fixes and a delimiter the tidier now reads.

Rule A9 no longer counts an author's surname as missing when the text layer split its small capital first letter into a token of its own. That shape is the largest single reason a page was refused over the last two extraction runs, and it cost the Chord paper the page holding the entries twenty five of its citations point at.

The tidier now reads a math delimiter that stands on a line of its own, which is how two papers write a display with the inline delimiters. Fourteen stray delimiters across six papers were going into the corpus as literal backslash-paren.

Both of these are in the tidier and the checker, so pages already read pick up the delimiter fix through papers extract --retidy, which asks no model.

What's Changed

  • Read a math delimiter that stands on a line of its own by @tamnd in #225
  • Stop a surname in small capitals counting as a missing word by @tamnd in #226

Full Changelog: v0.8.15...v0.8.16