Write the seventh article: the verdict that was never given - #50
Merged
Conversation
Bilingual, both HTML editions generated, registered in PUBLICACION.md with social copy for X and LinkedIn in both languages. The story is #32 and it is the least flattering one this project has published, which is why it is worth writing: the exported report — the document a teacher prints and carries to an integrity meeting — had never once printed a verdict, for any document, in any language, since the feature shipped. The CLI called the same text "Strong signs of AI writing" in the same run. One engine, two answers, 340 tests, and not one of them compared the surfaces. Underneath it, the thing that actually stings: the project publishes a measured threshold of 25 and the product drew its line at 20, in prose, five points below its own calibration page and below the highest-scoring human text in the corpus. That is the exact failure this tool exists to point out in other detectors, sitting in public next to the measurement. Also carries what the committee contributed, since that is the part other people can reuse: both reviewers found the always-false condition independently, then disagreed about the fix, and the disagreement is what produced the distinction between borrowing an aggregate *rate* (forbidden — it misstates reliability for one language) and an aggregate *boundary* (fine — it decides when the tool speaks, and the language's own figure is printed beside it). Verified against the CLI, as every article here is: 8/100 English, 7/100 Spanish, both below the measured boundary. The first draft was flagged for em-dash overuse — 20 in 1,749 words, the rhythm crutch the article itself describes — so half of them are gone and the finding with them. An article about a tool that contradicted itself should not ignore what the tool says about the article. Cover art still to draw; front matter points at the file name the other six use.
One document at the top, forking into two panels: the same score, and underneath it one solid bar where the command line stated a verdict against a dashed, struck-through one where the report gave none. That is the article's first paragraph with no paragraph. Wordless on purpose, like the citations cover. The only marks are the numbers and a ≠, which read the same in English and Spanish, so one file serves both editions. The verdicts are bars rather than text for that reason: the moment the cover sets "Strong signs of AI writing" it stops being bilingual, and this is a project that argues about exactly that. The first draft put the ≠ level with the two scores, where it read as "90 ≠ 90" and looked like an arithmetic mistake. The scores agree — that is the whole point — so the glyph moved down to sit between the two verdict lines, which are what disagree. Rendered with headless Chrome rather than ImageMagick: on this machine `convert` is the Windows filesystem converter, not ImageMagick, and it should not be handed a path and a density flag.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Seventh article, bilingual, both HTML editions generated, registered in
PUBLICACION.mdwith X andLinkedIn copy in both languages.
Documents #32 (PR #48) and the teacher-package correction (PR #49).
Why this one is worth publishing
It is the least flattering story this project has told:
never once printed a verdict, for any document, in any language, since the feature shipped.
The CLI called the same text "Strong signs of AI writing" in the same run.
40 where the rest cut at 45).
20, five points below its own calibration page and below the highest-scoring human text in the
corpus. That is precisely the failure this tool exists to point out in other detectors.
It also carries the reusable part — the committee found the always-false condition independently on
both sides, then disagreed about the fix, and that disagreement produced the distinction between
borrowing an aggregate rate (forbidden: it misstates reliability for one language) and an
aggregate boundary (fine: it decides when the tool speaks, with the language's own figure printed
beside it).
Verified
Run through the CLI like every article here:
The first draft was flagged for em-dash overuse — 20 in 1,749 words, the exact rhythm crutch the
article describes. Half are gone and the finding with them. An article about a tool that contradicted
itself should not ignore what the tool says about the article.
Still to do
Cover art (
social/verdict-story-cover.{svg,png}) — front matter already points at the name theother six use.