Releases: chethan62/grammar-server
Release list
grammar-server v0.4.3
grammar-server v0.4.3
Rule ids in every response are now LanguageTool rule ids. Harper names several
rules after the words they catch — The, Cant, OpenCompounds (a lot),
RepeatedWords, ItsContraction — and unmapped ones went out as rule.id
untouched. A client that filters, groups or displays by rule id saw "The" where the
contract promises a rule name: invisible to LTeX and anything else matching on
LanguageTool ids, and it reads as a word rather than a rule. Thirteen of them were
reaching responses on ordinary prose; every one now presents the id, category and
issueType that api.languagetool.org/v2/check reports for the same mistake, measured
rather than guessed (IT_IS, MODAL_OF, ENGLISH_WORD_REPEAT_RULE, TOO_DETERMINER,
CANT, WERE_WHERE, WONT_CONTRACTION, COMMA_PARENTHESIS_WHITESPACE, THERE_THEIR,
and MORFOLOGIK_RULE_EN_US for the pure typos).
enabledOnly and enabledRules accept either spelling as before — the filter compares
the LanguageTool id and harper's native name — and the two new category ids (MISC,
CONFUSED_WORDS) are in the set /status-visible checks can actually emit, so naming
them is not reported as an incomplete check.
Two tests keep it that way: every id in the mapping table must be LanguageTool-shaped and
carry a category the server can emit, and a corpus of the sentences that used to leak is
driven through the real engine and asserted to come back with LanguageTool-shaped ids
only. Removing one mapping turns both red.
Also in this release: the bundle carries grammar-ui v0.4.4 (rephrase candidates can no
longer splice a stale range, and the Fix-sentence icon is an SVG instead of a glyph).
v0.4.2 — one sentence boundary; fix-sentence says what it fixed
One sentence boundary, and a fix that says what it fixed.
Fixed
- The sentence boundary existed twice — a period scan in the API and the
abbreviation-aware one the stats use — so they disagreed on every title and initial:
"R. K. Rao go to Mysore."was treated as"Rao go to Mysore."and/v2/fix-sentence
fixed that fragment. Both come fromlt.SentenceRangesnow. /v2/fix-sentenceanswers{fixed, offset, length}: the range it fixed, in UTF-16 code
units, so a client replaces exactly that text. grammar-ui uses it (without it, a
fragment replaced by a whole sentence duplicates text:"Dr. Dr. Smith…").- Spelling suggestions are no longer applied to names: harper reads
"R."as a typo
(→ "RI") and"Rao"as one (→ "Rad"). A single letter or a capitalised word that is
not the first in the text is left alone; grammar and typography suggestions still apply,
soShe go to the office.→She goes to the office. - The contract is a CI job now (
contract), not a thing someone remembers to run: it
builds, starts the server, waits for/status, and runsexamples/lt-client-smoke.py
againstlanguage_tool_python. The harper pin moved into one composite action, since
two jobs needing it is how a pin drifts.
Verified: full suite, gofmt, vet; live probes for each fix-sentence case; the browser
path ("Dr. Smith wrote teh report. It was late." → one click → "Dr. Smith wrote the report. It was late.", no duplication, the second sentence untouched); CI green on both
jobs.
This archive: grammar-server, harper-ls, harper-cli (the server finds the pair in
its own directory, so nothing needs installing and PATH is irrelevant), the systemd unit,
the UI's three files with their unit, README.md, LICENSE, LICENSE-harper. Extract
and run ./grammar-server.
v0.4.1 — what exercising the shipped product found
Delivery fixes found by exercising the shipped thing, not the source.
Fixed
- Chunk boundaries fell inside words on text with no
.,!or?(bullet lists,
tables, comma run-ons): the engine saw half-words and reported both halves as
misspellings. 29 KB of clean text came back as'o'@11998+'ver'@11999; cuts now
land on the last space, tab or newline. /v2/statscounted every period as a sentence end, soDr. Smith,1.5 lakh,
R. K. Narayan,10.30 a.m.,example.comand...all inflated the sentence
count — and mean/longest sentence length, Flesch ease, FK grade and the grade band
are computed from it./v2/fix-sentenceread the client's UTF-16 offset as a byte index: with an emoji or
an accent before the error it fixed the previous sentence, and grammar-ui swapped
that sentence into the document.enabledOnlynaming rules this engine does not have answered "no issues" — a clean
document that was never checked.warnings.incompleteResultsis now true instead.- The binary reported a hardcoded
0.4.0regardless of the revision it was built from;
the version is now injected from the git tag, so/statussays0.4.1(or
0.4.1-3-gabc1234for commits on top of it). - The release archive could carry the previous release's binary (tagging changes no
source file, and the build target only watched sources) — the build now always runs.
Measured on this build (66 words at level=picky, 2 findings): p50 15 ms / p95 16 ms
with one request in flight, p95 91 ms at 4 in flight, ~65 checks/s at 32 in flight, no
failures. Checks are serialized by design, so throughput is flat and latency queues.
Verified against a real client: language_tool_python passes five contract
assertions (typos, correct(), the picky style tier, disabledRules, enabledOnly) —
uv run --with language_tool_python python examples/lt-client-smoke.py. Hostile input
(every method × every endpoint, 30 absurd bodies each, 6 content types) answers 4xx and
never 5xx, and never wedges the engine.
This archive: grammar-server, harper-ls and harper-cli side by side (the server
finds the pair in its own directory, so nothing needs installing and PATH is irrelevant),
the systemd unit, the UI's three files with their unit, README.md, LICENSE and
LICENSE-harper (harper is Apache-2.0). Extract and run ./grammar-server.
v0.4.0 — LanguageTool contract, style tier, make install
Offline grammar, style and delivery checking behind the LanguageTool /v2/check API — no Java, no cloud, no account. harper-ls is the engine.
Request contract — everything LanguageTool clients send: enabledRules/disabledRules (LanguageTool rule ids or harper's own names), enabledCategories/disabledCategories, enabledOnly, level, preferredVariants, motherTongue, data=/text, warnings.incompleteResults. Rule toggles reach the engine, so rules harper ships off can be switched on.
Style tier (opt-in via level: "picky" or enabledCategories: ["STYLE"]) — three deterministic passes, no model involved: WORDINESS (33 wordy phrases → the concise form), PREFERRED_TERM (10 non-preferred forms → the house-style one: e-mail → email, whilst → while), PASSIVE_VOICE_SIMPLE (LanguageTool's own id).
Engine — harper 2.11.0, 892 rules, 12 off by default. One persistent harper-ls, documents closed after each check, a quiet engine reconnects and retries once, and long documents are chunked (~12 KB, sentence-aligned) so offsets stay correct into the hundreds of kilobytes.
Also — GET /v2/stats (words, sentences, grade level, reading time), POST /v2/rewrite (optional, needs a local Ollama), GET /v2/languages, GET /status (reports the dialect the engine is on). Loopback by default; 200,000-character cap with LanguageTool's own 413 shape.
Install
make install # binary → ~/.local/bin, user unit installed
systemctl --user enable --now grammar-serveror make bundle-local for a portable directory to copy anywhere. Needs harper-ls on PATH (sudo pacman -S harper, brew install harper, or point --harper at it).
Measured on this machine — 65-word documents at level=picky: p50 62 ms / p95 102 ms with one request in flight (~20 ms of linter plus ~5 ms per match fetching suggestions), ~20 checks/s aggregate, zero failures or timeouts in a 183-request concurrent load test.
Verified — go test ./... green; CI runs gofmt, go vet and the full suite against harper 2.11.0.