Releases: HilbertraumAI/HilbertRaum
Release list
v0.1.62
Which file do I need?
The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:
| Your computer | Download this ONE file |
|---|---|
| Windows 10/11 | HilbertRaum-0.1.62-portable.exe |
| macOS (Apple Silicon) | HilbertRaum-0.1.62-mac-arm64.app.zip |
| Linux | HilbertRaum-0.1.62.AppImage |
The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.
This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Two optional downloads are offered in the app only when you need them: the knowledge-pack tools, and the text-recognition (OCR) files for scanned PDFs and photos (German + English, about 4 MB; offered on the scan or photo in Documents and on the AI Model screen). Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
Added
- Text recognition (OCR) can now be added from inside the app. When a scanned PDF or a
photo of a page needs the OCR language files and they are not on the drive, its row in
Documents offers Download OCR files — and the AI Model screen shows the same
offer as "Text recognition for scans and photos (optional)". A short confirmation names the
two files (German and English, about 4 MB, Apache-2.0) and where they come from; once they
are downloaded and checked, text recognition works without restarting the app: the scan
offers Make searchable (OCR), and a photo that failed reads with Try again. Like
every other download it asks first, needs the drive policy and Allow internet access… to
permit it, and is verified before use. Portable-app users no longer need the drive-setup
scripts for OCR; they remain the offline alternative (#410). - While you dictate, the app tells you at once when nothing is reaching the microphone.
After about two seconds without any sound, a note appears under the message box and
disappears as soon as sound comes in — so a muted or wrong microphone is caught before you
stop, not after (#497).
Changed
- The AI engine is updated to llama.cpp b11146 (the upstream v0.5.0 release). In our
document-question tests the ten built-in chat models answer as accurately as before. On the
Qwen3.5, Qwen3.6, Qwen3.8 and Gemma 4 models, coming back to a conversation after something else used the model (a
document being indexed, another conversation) no longer means re-reading the whole
conversation first: the engine now restores it from memory, so that first reply starts sooner.
The memory the engine keeps for this is limited to an eighth of the computer's RAM, at most
8 GB; when it is full, the oldest saved conversation is dropped. A drive set up with the
previous engine keeps using it until the drive-setup script is run again (#512). - Document translation works in slightly smaller parts. A long document is now translated
in parts of about 640 words instead of about 690, so a part full of dense text, such as an
invoice or a table, stays within the input size the translation model was trained on. A long
document takes about 7 % more parts (#512). - Dictation now ignores background noise. Before transcribing what you dictated, the app
runs a small voice-activity model (Silero VAD) so a recording with no speech in it yields
"No speech was recognized" instead of a stray word. The model is a second, small file of the
speech model: the AI Model screen fetches it together with the speech model, and a drive
that has only the older single file shows the speech model as incomplete until that one
file is added (#504). - The microphone button no longer disappears when the speech model is missing. On a
drive without the speech model, the chat composer now shows the button greyed out; clicking
it explains what is missing and opens the AI Model screen, where the model can be added
(#497).
Fixed
- A model that cannot be loaded is now reliably told apart from a graphics-card problem.
Before, on some systems such a model could switch graphics acceleration off for every model,
or two different start failures could be taken for the same one (#515). - Inline math written as
$…$now renders in chat. Before, only$$…$$blocks and
\(…\)spans were typeset, so the inline formulas many local models write showed up as raw
text with the dollar signs. A single-dollar span is typeset only when it looks like math, so
amounts such as "$5 and $10" still read as plain text (#501). - The "reply cut off" note no longer blames the context window for every cut. A grounded
answer from your documents or a knowledge pack is stopped by a fixed length the app allows
for one reply, and the note used to say the model's context limit was reached and to suggest
raising the context size, which cannot help there. The note now just says the reply was cut
off, and its tooltip suggests raising the context size only when the context really ran out
(#498). - Superscripts and subscripts in knowledge-pack articles stay readable everywhere. An area
of25 m²or a number like10⁶in an ordinary paragraph used to reach the model and the
article view fused as25 m2and106; only table text kept them apart. Paragraphs and
headings now use the same readable form as tables (25 m^2,10^6,H_2O), and search
still finds them from a plainly typed question such as "CO2" or "106" (#488). - Voice dictation works right after you install the speech model or the voice engine — no
restart. The app used to keep its startup answer ("not available") until it was restarted,
without saying so; the microphone button now appears as soon as the download or install
finishes, also when you stayed in the chat meanwhile (#497). - A silent recording no longer turns into a stray word. When nothing reaches the
microphone — it is muted, the wrong input device is selected, or the operating system does
not allow the app to use it — dictation now says so instead of inserting the word "you"
(#497). - Two small article-reading fixes: a rare text leak, and photo captions. Reading a
Wikipedia table that contained a certain very rare block of raw text could leak a stray
fragment of it into what the app reads back; that block is now skipped correctly, the same
way it already is outside of tables. And a photo's caption, which used to be dropped along
with the photo itself, is now kept as part of the article's text — the app reads captions on
purpose now, the way it already reads a table's own caption.
Issues resolved
- #410 — OCR language files are the only drive asset with no in-app install path (follow-up to #59)
- #436 — Accessibility: ErrorBanner's failure message is never announced — the nested inner role="status" defeats the always-mounted alert (M-U1)
- #460 — Tests: 99 of 129 openDatabase() suites never close the handle — ~1,450 locked temp roots per Windows run
- #488 — Superscripts and subscripts are readable in table text but still fuse in prose
- #493 — A CDATA section inside a kept table can still leak its text, where ordinary paragraphs do not
- #498 — Cut-off badge and its remedy hint blame the context window for the app's own cap
- #501 — Chat: inline math written as
$...$ shows as raw text, only$$...$$ blocks render - #504 — Silero VAD for whisper: fetched in-app with the speech model, used for dictation (follow-up to #497)
- #512 — Bump the llama.cpp runtime pin b9849 → b11146 (v0.5.0)
- #515 — Start ladder: the #312 model-vs-device check compares non-comparable log lines (colour reset / per-process timestamps)
- #517 — Grounded-QA scorer misses German abstentions ("nennt … nicht", "nicht festgelegt", "nicht direkt genannt")
- #518 — Re-capture the --list-devices and vision SSE fixtures on the current pin (F-40)
What's Changed
- fix(zim): CDATA in kept tables, and figure captions (#493) by @comilionas in #500
- fix(dictation): restart-free speech-model activation, silence gate before whisper, visible "not installed" mic (#497) by @comilionas in #503
- feat(chat): live "no sound is reaching the microphone" hint while dictating (#497 follow-up) by @comilionas in #505
- feat(dictation): Silero VAD before whisper for dictation, fetched in-app with the speech model (#504) by @comilionas in #506
- fix: readable sup/sub in knowledge-pack prose (#488), cause-neutral cut-off badge (#498), inline
$…$ math (#501) by @comilionas in #509 - OCR language files install in-app and activate without a restart (#410) by @comilionas in #510
- Translation planner constants 2.8 in / 3.1 out (#512 decision 3) by @comilionas in #513
- llama.cpp b...
v0.1.61
Which file do I need?
The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:
| Your computer | Download this ONE file |
|---|---|
| Windows 10/11 | HilbertRaum-0.1.61-portable.exe |
| macOS (Apple Silicon) | HilbertRaum-0.1.61-mac-arm64.app.zip |
| Linux | HilbertRaum-0.1.61.AppImage |
The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.
This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
Added
- Knowledge-pack answers can now draw on the values inside a page's tables. Material property
tables, comparison tables and similar structured content used to be dropped entirely when a
knowledge-pack page was read; those values now reach the pack's content and can be cited in an
answer, the same way ordinary paragraph text already was. - A select-all checkbox in the documents view. Next to the per-row checkboxes and the "Add to
project" action, one control now selects or clears every document in the current list at once
(#214).
Changed
- Grounded document and knowledge-pack answers no longer vary in wording between two askings of
the same question. The app now sends the model a fixed decoding setting (no randomness, and a
per-pass length cap that is continued rather than cut off) instead of the model server's own
defaults, so a repeated question gets the same wording back. A related case, where the cut-off
notice itself can be inaccurate, is tracked separately (#498). - Knowledge-pack answers only name a pack's language when it is relevant. The search planner
used to mention the pack's archive language on every question; it now does so only when the
question is asked in a different language than the pack itself. On a maintainer benchmark this
roughly doubled how often an English question against a German-language pack reached the right
article (27 in 100 to 50 in 100), with German-language questions unaffected (#486).
Fixed
- Knowledge-pack table content no longer leaks a page's stylesheet or repeats a formula. A
style/script element or a mathematical formula nested inside a table that a knowledge pack reads
could end up in the delivered text — raw CSS in one case, both the rendered and the source form
of the same formula in the other. Both are now handled the way the rest of the page already was:
styling is dropped, and a formula is converted once, not twice (#485, #490). - The reranker recovers from a graphics-card failure instead of staying off for the session.
When ranking on the graphics card fails, the app now falls back to the processor for the rest of
the session, and the request-size ceiling that fallback enforces is correctly re-applied along
every path that can trigger it (#474). - The document-search model's shutdown can no longer overlap with itself. (#475)
- The Performance page now accounts for a loaded ranking model. Its graphics-memory figures
used to check only the chat and translation models; they now include a resident ranking model
and the graphics memory it is contending for (#476). - The counter behind a pending model switch can no longer silently read as empty. (#477)
- A test covering evidence-pack PDF text extraction no longer misreports a passing result as a
failure. (#481) - The release build now runs its test suite split into shards, the same way ordinary CI already
does, instead of one long unsharded run. (#458)
Issues resolved
- #214 — Add a "Select all" checkbox in document upload view
- #447 — The ZIM query expander was never measured against the #399 prompt-cache eviction (inference, not measurement)
- #458 — CI: both Windows legs run at their time budget and flake — three re-run cycles in two days
- #474 — Reranker recovery after a GPU failure, including a GPU→CPU demotion
- #475 — The E5 embedder's teardown overlap across workspace lock and quit (pre-existing)
- #476 — The Performance card summary line omits a GPU-resident reranker
- #477 — Make the reranker's pending-model-switch counter impossible to leave unwired
- #478 — Deliver tables to the model instead of dropping them
- #481 — evidence-pack-pdf-smoke: the two archive-pack cases fail on master where the font splits the "fi" ligature (skipped in CI)
- #485 — Stylesheet source can leak into delivered table text and into citation snippets
- #486 — English questions retrieve far less than German ones from a German Wikipedia pack
- #490 — Formulas inside tables are delivered twice, as loose markup characters plus raw TeX source
What's Changed
- eval(#447): the ZIM query expander evicts the prefix on every pack-scoped turn — measured by @humaniser in #480
- feat(zim): deliver tables to the model instead of dropping them by @comilionas in #479
- docs: committed eval evidence under eval/results/ is not "logs" in the hard rule by @comilionas in #484
- docs: record the knowledge-pack retrieval research outcome and its negative result by @comilionas in #489
- fix(zim-tables): stop a style/script body and a doubled math formula from leaking into kept table text by @comilionas in #492
- fix(rag): pin grounded answers to a fixed decoding setting by @comilionas in #491
- UI select-all, PDF-smoke ligature fix, and release.yml Windows sharding by @comilionas in #494
- Reranker recovery, embedder teardown, Performance card, and a required-wiring fix by @comilionas in #495
- zim: plan in a pack's language only when the question's differs (#486) by @comilionas in #496
- docs: release notes, shipping rule and BUILD_STATE drain before the next tag by @comilionas in #499
- release: v0.1.61 — version bump by @comilionas in #502
Full Changelog: v0.1.60...v0.1.61
v0.1.60
Which file do I need?
The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:
| Your computer | Download this ONE file |
|---|---|
| Windows 10/11 | HilbertRaum-0.1.60-portable.exe |
| macOS (Apple Silicon) | HilbertRaum-0.1.60-mac-arm64.app.zip |
| Linux | HilbertRaum-0.1.60.AppImage |
The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.
This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
Added
- Knowledge packs, one click from the Documents title. The Documents screen now has a switch
next to its title — My documents | Knowledge packs — instead of a "Reference" entry at the
bottom of the sidebar. In the packs view each pack has an Enabled switch, an Ask this
pack button that opens a chat answering from that pack alone, and a ⋯ menu for Remove.
The first-run notice became a setup card, and the empty view nameslibrary.kiwix.orgwith a
Copy the library address button. Home's readiness card gained a Knowledge packs row,
and the chat's "Answering from" picker offers Add packs… while you have none. - Filter documents by name. A search box above the document list narrows the current
section as you type. - Check all model files, and stop it when you want to. A new action near the top of the AI
Model screen checks every model file on the drive against its published checksum. It replaces a
check the app used to run on its own every time you opened that screen (see Changed). On a slow
drive the full check can take several minutes, so Stop checking sits beside its progress bar:
the files it already checked stay checked, the rest are simply left unchecked, and stopping is
not treated as a failure. The check keeps running if you move to another screen — come back and
the progress bar and Stop checking are still there. - A quieter left navigation. The sidebar now has three groups: Chat, Documents, Translate
and Images for everyday work; AI Model and Performance for the machine; Settings at the
bottom. The HilbertRaum mark at the top is the Home button and lights up when you are on
Home, so the separate Home entry is gone. Skills moved into Settings as its own tab; the
skill picker in the chat composer is unchanged. - A new Performance page in the left navigation. It answers, in plain words, what this
computer can run and how fast: one sentence, then four tiles for speed, memory, graphics
memory and drive, each with a rating word, and a "Your model" line that says where your
current model lands on this computer: in graphics memory, partly on the graphics card with
the rest in RAM, on the processor, or too large, with the sizes behind the verdict once the
model has started. Apple Silicon shows one unified memory figure. A "Models on this computer"
card lists every model the app can hold (chat, translation, images, document search, voice),
where each runs and whether it is loaded right now, and sums what the graphics card and the
RAM would need with everything loaded at once. Below it, figures from real use (your last answer, the last model
start, the last file check) and one result for every computer this drive has been plugged
into. The check runs on its own the first time and whenever the drive lands on a different
computer; on a computer it already knows, the earlier result is restored so the recommended
model follows the machine. While a check runs you see its steps instead of a plain "Running…"
button, and if no model has run yet the page offers to start the recommended one and measure.
The technical table stays on Settings → Diagnostics. English and German. - A new computer's first check never borrows another computer's drive-speed reading.
Plugging this drive into a computer it hasn't seen before measures that computer's own read
speed from scratch; a warning about a slow drive, or the lack of one, always reflects this
computer, never whichever computer the drive was in last. - The Performance page refreshes itself and stays precise about what it measured. It updates
in place the moment a check finishes, a model starts, or a file is verified, with no need to
leave and come back. A speed reading counted from streamed chunks rather than the model's own
timing is marked Approximate everywhere it appears, including another computer's row and the
copied report, which also always names the computer it describes. The graphics-memory tile
says plainly when a chip's memory is Integrated and shared with the rest of the computer, and
shows Not recorded, instead of guessing, for a computer this drive visited before the check
could record its card. When a model only partly fits in graphics memory, the explanation now
states the exact safety margin the AI engine reserves. After a check you started with the
keyboard, focus returns to Check again. - Checking your computer's speed no longer competes with starting a model. The automatic
first-time check now waits for a model that is already starting up to finish loading before it
measures your drive, instead of reading and loading at the same time. - Knowledge packs: ask an offline Wikipedia (or any ZIM archive). Register ZIM
files — e.g. from the Kiwix library — as knowledge packs (Documents → Knowledge
packs, or just drop them into the drive’szim/folder), tick them as sources in a
documents chat, and answers draw on them with citations that name the archive and
open the article offline. Fully local: the pack server binds to 127.0.0.1 only,
archives are used in place and never copied. Needs the kiwix-tools binaries on the
drive — still a manual step in this release (see the user guide §7b). - Keep a knowledge-pack article in your documents. Save to my documents files a copy of a
cited article alongside everything else you have imported, so it stays searchable even after
you disable or remove that pack, or unplug the drive it lives on. The action sits on the
citation itself, next to Open article, and inside the article view; it names the copy it
filed, and saving the same article twice tells you it is already there instead of making a
second one. Reviews stay read-only, so it is not offered there. - Knowledge packs are found once, not on every open. The list is discovered when you
unlock and by an explicit Refresh in Documents → Knowledge packs, instead of a
fresh drive scan every time the panel opens or a message is sent — packs load
instantly, and copying a new archive onto the drive shows up after Refresh. A pack
whose file was replaced by a different archive is now shown as such ("Different
archive") instead of quietly answering from the wrong one. Locking or quitting the app
stops the pack server and removes its small generated index file. Opening an alias article
(a redirect entry — about half of a Wikipedia archive's titles) shows the article it points
to instead of an error. A large article that the bundled Windows pack server occasionally
delivers only in part is re-requested instead of being reported as unavailable or dropped from
the answer. Known limit on Windows: an archive can only be added from a path
without umlauts or accents (the bundled kiwix-manage cannot read such paths) — the drive's
zim/folder always works. - Evidence reviews now name the knowledge pack, not a same-named document. Reviewing an
answer that cites a knowledge-pack article records the archive, the article and its pack
id honestly, and shows identity as not verifiable against the workspace instead of
matching it to a similarly named document — the review, its HTML/PDF evidence pack and
the Markdown transcript export all name the pack. Reviews created on a pre-release
knowledge-pack build that cite an archive must be re-run. - Adding several knowledge packs at once now reports exactly what happened. If some of
the chosen archives could not be added, you are told how many were added and how many
were not — never the technical reason a file failed. Answers and the article viewer now
double-check with the pack server that it is still the same one that was running a moment
ago, and retry once if it was restarted mid-question, instead of silently trusting
whatever answers on that port. One privacy limit to know: other programs running under
your own user account can read the enabled packs while the workspace is unlocked, through
the pack server, which has no password of its own; locking or quitting the app stops it. - Answer from knowledge packs alone, and see what each one did. A new Search my
documents toggle in the sources picker lets you answer from ticked knowledge packs
only — files attached to the chat are still used either way — and every answer now
shows a "Knowledge packs:" line ...
v0.1.59
Which file do I need?
The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:
| Your computer | Download this ONE file |
|---|---|
| Windows 10/11 | HilbertRaum-0.1.59-portable.exe |
| macOS (Apple Silicon) | HilbertRaum-0.1.59-mac-arm64.app.zip |
| Linux | HilbertRaum-0.1.59.AppImage |
The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.
This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
Fixed
- A second running copy of the app can no longer destroy your encrypted workspace
(issue #208). Starting the app while another copy was already running — the natural
upgrade flow: launch the new version, then close the old one — could silently and
permanently corrupt an encrypted workspace on Windows, and the damage only surfaced at
the next unlock, looking like a wrong password. Three fixes ship together: the app now
refuses to start a second copy (the running copy's window comes to the front instead);
the workspace re-encryption on lock/quit refuses to overwrite the vault with anything
that is not actually your database; and a workspace whose encrypted data is damaged now
says so plainly at unlock — with backup guidance — instead of implying the password was
wrong. See the new troubleshooting entry "Your password is correct, but the workspace
data on the drive is damaged". - Failing to open the workspace database no longer leaves the decrypted file behind on the
drive, and no longer holds an invisible open file handle to it.
Changed
- Developer runs (
npm run dev) now use their own app-data folder (…\@hilbertraum\ desktop-dev) instead of sharing the released app's production workspace. A workspace
previously created from a dev run stays on disk under the old path; point the dev run at
it withHILBERTRAUM_DRIVE_ROOTif you still need it.
Issues resolved
- #208 — Encrypted workspace destroyed on disk: .enc now decrypts (with a valid tag) to random noise, and the unlock failure looks like a wrong password
What's Changed
- docs(readme): rework README for non-technical users by @humaniser in #209
- fix(vault): a second running instance could destroy the encrypted workspace (#208) by @comilionas in #210
- release: v0.1.59 — version bump by @comilionas in #211
Full Changelog: v0.1.58...v0.1.59
v0.1.58
Which file do I need?
The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:
| Your computer | Download this ONE file |
|---|---|
| Windows 10/11 | HilbertRaum-0.1.58-portable.exe |
| macOS (Apple Silicon) | HilbertRaum-0.1.58-mac-arm64.app.zip |
| Linux | HilbertRaum-0.1.58.AppImage |
The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.
This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
Added
- Local API (optional, off by default). Other programs on the same computer can be
allowed to use the running chat model over a loopback-only, OpenAI-compatible endpoint
(http://127.0.0.1:4980/v1—GET /v1/modelsandPOST /v1/chat/completions, streaming
or not, with JSON-schema-constrained output). It is off until switched on behind a consent
dialog, requires an access key by default, never reaches the internet, exposes no documents
or conversations, keeps no record of what was asked or answered, exists only while the
workspace is unlocked, gives your own chat priority over any outside caller, and can be
forbidden by drive policy. Tutorial, client examples and the wire contract:
docs/local-api.md. - Faster answers from the large Qwen3.8 models on a capable GPU (issue #182). These model
files always carried a small built-in "draft" head that the engine loaded and ignored; it is
now used, worth a measured 38–45 % faster text generation on the reference machine. It
engages only when one GPU has room for the model plus ~3.5 GiB, and any refusal falls back
to exactly the previous behaviour — a model can be slower to start once, never broken.
Changed
- New recommended chat models for 24 GB and ≥32 GB machines (issue #196). Unsloth removed
the Qwen3.8 files the previous recommendations pinned, so those downloads had begun failing
with a 404. Two things changed. A model whose upstream source has been withdrawn now says
so on the AI Model screen, with the reason and the date, instead of offering a Download
button that cannot work — and the drive-provisioning scripts skip it with a clear line
instead of retrying a dead link. And the closest published successors were re-measured on
the reference machine before taking over the tiers: Qwen3.8 27B UD-Q4_K_M at 24 GB and
UD-Q5_K_M at ≥32 GB. Answer quality is unchanged within measurement noise; the 24 GB
pick generates about 19 % slower than the withdrawn file, which is recorded in the
catalog rather than glossed over. Already-downloaded copies of the old files keep
working — they still verify, still start, and keep the speed-up above.
Fixed
- Downloading "Gemma 4 12B Instruct QAT Q4" works again (issue #201). Google replaced that
model file on their servers in July with a corrected version, at the same address. The app
checks every download against the exact file it expects, so it was transferring the full ~7 GB
and only then reporting a checksum failure — with nothing you could do about it. The catalog
now points at the corrected file, which was measured against the old one first and answers
identically. If you already downloaded this model, your copy will now report a checksum
failure and needs downloading again (~7 GB) — press Download on the AI Model screen and it
replaces itself. No other model in the catalog is affected: all the rest were re-checked
against their sources in the same pass. - Deleting a document really deletes its stored copy after a drive changes letter
(issue #188). The workspace recorded each imported copy by absolute path, so moving the
drive between computers — or just getting a different drive letter — left every stored copy
unreachable: "Delete document" removed the entry while the encrypted copy stayed on disk,
and exporting or previewing the original failed. Paths are now resolved relative to the
drive, and existing entries heal themselves the first time each document is opened. - Re-indexing or deleting a single document now tells you what happened (issue #194).
The action worked but reported nothing at all — no confirmation, no progress, no error. It
now shows a spinner on that row, confirms with the document's name, and names the document
in the message if it fails. "Re-index all" already behaved this way. - One AI job at a time, across every part of the app (issues #185, #186). Starting the
Diagnostics benchmark, or a skill run, could put a second job on the model while an answer
was still streaming — each part of the app tracked "busy" separately and none of them
agreed. They now share one signal. A speed measurement taken while something else was
running is discarded rather than saved, so a contended reading can no longer push your
machine's model recommendation down permanently; when that happens, the benchmark says so.
Security
- Electron 39.8.10 → 43.4.0 (issue #179), clearing the last open dependency alert:
GHSA-jmr9-qjv8-65gv (CVE-2026-56876, high) inextract-zip, which is abandoned at its last
version and could only be cleared by moving Electron. Chromium 142 → 150, Node 22 → 24,
SQLite 3.51.2 → 3.53.1. Minimum operating systems are unchanged (Windows 10+, macOS 12+)
and no stored format changed, so existing workspaces open exactly as before.npm audit
reports 0 vulnerabilities.
Issues resolved
- #179 — Bump Electron 39 → 43 (removes the unpatchable extract-zip dependency, Dependabot #83)
- #182 — Enable MTP speculative decoding for the Qwen3.8 chat manifests (+38-45% measured decode)
- #185 — Benchmark has no re-entrancy or busy guard (surfaced by the local-API generation gate)
- #186 — A skill run and a chat stream can generate concurrently (surfaced by the local-API generation gate)
- #188 — Document export offers "Originaldatei exportieren" for documents it cannot export — and the stored copy may only look missing (absolute stored_path)
- #194 — Single-document re-index gives no success or progress feedback, while "Re-index all" toasts and shows a progress bar
- #196 — qwen3.8 manifests: upstream deleted the pinned GGUFs (Dynamic 3.0 restructure), all three download URLs return 404
- #201 — gemma4-12b-it-qat-q4: upstream re-uploaded the GGUF (corrected vocabulary) — the pinned SHA-256 no longer matches, so downloads fail verification
- #202 — Four chat manifests carry rounded, estimated download.size_bytes instead of the real byte count
What's Changed
- docs(build-state): archive closed-wave entries (retention ritual) by @comilionas in #183
- Local API endpoint wave (P1-P6) by @comilionas in #184
- docs: §9.4 addendum — measured MTP findings for Qwen3.8 (not adopted) by @humaniser in #181
- Electron 39 → 43: clear Dependabot #83 (unpatchable extract-zip) by @comilionas in #187
- fix(ingestion): portable stored-copy resolver — "delete" now actually deletes (#188) by @comilionas in #189
- fix(runtime): one shared occupancy span for the in-app model lanes (#185, #186) by @comilionas in #192
- feat(runtime): MTP speculative decoding for the Qwen3.8 Q4/Q5 chat manifests (#182) by @comilionas in #191
- fix(tests): the read-only stored-copy diagnostic — rebuilt, run, and its password path fixed (#190 phases 1–2) by @comilionas in #193
- fix(docs-ui): single-document re-index/delete now report success, progress and a named failure (#194) by @comilionas in #195
- docs(local-api): document the local API as a first-class release feature (O2 lifted) by @comilionas in #197
- docs(build-state): record where the #196 fix actually landed (O6 citation correction) by @comilionas in #198
- #196 successor wave: measured UD manifests, ranks ratified (generational handover restored) by @humaniser in #199
- docs(changelog): per-version release notes — the file is the release page, so it has to be true by @comilionas in #200
- docs: 2026-08-20 audit remediation — the #196 successor wave never reached the catalog prose by @comilionas in #203
- docs(build-state): 1,968 -> 763 lines — the retention rule was draining the wrong section by @comilionas in #204
- fix(catalog): re...
v0.1.57
Which file do I need?
The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:
| Your computer | Download this ONE file |
|---|---|
| Windows 10/11 | HilbertRaum-0.1.57-portable.exe |
| macOS (Apple Silicon) | HilbertRaum-0.1.57-mac-arm64.app.zip |
| Linux | HilbertRaum-0.1.57.AppImage |
The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.
This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.
Added
- Local chat — a
llama.cppruntime running GGUF models entirely on-device
(CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
Granite) and an on-machine benchmark that recommends the best-fit model for
available RAM. A built-in demo mode runs the whole UI with no model files
and no network. - Document Q&A with citations — import PDF / Word / text, ask questions, and
get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
reranker, scoped by collections. - Image understanding — ask questions about a picture with a local vision
model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
and deletable — per-entry and total sizes are shown, a confirmed "Clear
history" removes everything at once (issue #122), and the history's disk
footprint appears in Settings → Diagnostics. - Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
on-device OCR ("Make searchable") on scanned pages (bundled German + English
language files; no cloud OCR). - Document tasks & skills — summarize, translate, and compare documents;
install reusable skills for structured extraction (bank statements,
invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe). - Evidence packs (review mode) — review a document-grounded answer block by
block against its sources, record explicit decisions and notes, and export the
review as a self-contained evidence pack in HTML or PDF — generated
locally and offline, with honest coverage, freshness, and limitation notes. A
pack supports human review; it is not a correctness certification. - Translate view + dedicated translation model (TranslateGemma) — a top-level
Translate screen for live text translation and drag-and-drop document
translation across 51 languages, source and target (the model's full
production tier — from German and English to Arabic, Chinese, Swahili, and
Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
sidecar (downloaded on demand behind the license acknowledgement; not
bundled) — never the chat model; document translations materialize as
searchable, exportable local documents. GPU-accelerated when the machine
allows it, with an automatic CPU fallback. Calibrated against the real model
(per-language round-trip evidence + measured tokenizer weights) so a window
can only over-chunk, never overflow. - Encrypted, portable workspace — an optional password-encrypted workspace
(AES-256-GCM with Argon2id key derivation) covering the database, imported-document
copies, and the diagnostics log; keep models plus the workspace on an external
drive and move between laptops. - Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
context-isolated renderer; a strict Content-Security-Policy; deny-by-default
renderer permissions; an offline guard that trips on any non-loopback connection
attempt; and a tamper-evident local audit log (ids/counts only, never content). - Cross-platform & distribution — Windows-first, with macOS and Linux supported
in the architecture; portable / preconfigured-drive distribution via
scripts/build-commercial-drive.*. - Standard project docs — this
CHANGELOG.mdand a
CODE_OF_CONDUCT.md. - Windows engine build attached to releases (issue #102 item 1). Each release
now carriesllama-runtime-win-x64.zipbeside the existing mac Metal zip: the
same SHA-256-verified llama.cpp build the in-app installer fetches, packaged so
it unzips straight into the drive'sruntime/llama.cpp/folder — an offline /
air-gapped install path that needs no repo scripts. - Opt-in perf mark log for measurement runs: with
HILBERTRAUM_PERF_LOG=1set in
the environment, the app appends timestamped timing marks (startup phases, vault
unlock/lock split, model checksum vs. load split, time to first token, document
ingestion phases) tologs/perf.log. Off by default; no file is created without the
variable. Records the cold-load figuresdocs/model-benchmarks.md§11.4 lists as
still missing. Seedocs/benchmark.md"Perf marks". - Model checksum passes are visible now (issue #106). Verifying a model's
multi-GB file — which can take minutes from a slow USB stick and used to run
with no trace anywhere — now writes one diagnostics-log line per real
verification (which model, how many bytes, how long), pluschecksum_start/
checksum_donemarks in the opt-in perf log. Overlapping verifications of the
same file (for example a start racing an AI Model screen visit) also no longer
each read the whole file — they share one pass. - Honest progress while a model starts (issue #107). Once the app has
measured your drive's real read speed, the Chat screen's "your model is
starting" panel shows the model file size and an approximate percentage read
instead of an indefinite spinner. Fresh installs (no measurement yet) keep the
plain message rather than showing a made-up number. - Cold model starts are much faster on slow drives (issue #114). While a
model loads, the app now also reads the model file front-to-back in the
background, which lines the operating system's file cache up with what the
load is about to need. Measured on a 16 GB laptop: a 6.65 GB model from a slow
USB stick started in 5:46 instead of 11:20 (−49%), and from a portable SSD in
10 s instead of 15 s (−36%). Starts that are already fast stay as they are —
the helper is skipped when the file was just verified (its content is already
in the cache) and stops the moment the load finishes or is cancelled. - One-click skill offers on document answers (issue #80). When a document
answer cannot serve the shape you asked for — "categorize the transactions
and sum per category" answered by an engine that can only list values — the
answer now carries a one-click "Run "Bank Statement Analysis" for this
question" action instead of only a prose pointer; clicking re-answers the
same question with the skill (nothing ever runs without the click). On the
residue the deterministic router provably cannot classify, a bounded,
grammar-constrained on-device model pick (marked "Suggested by the local
model") may fill the same offer; ordinary questions keep their zero-model-call
routing, and every failure quietly falls back to today's answer. - Export your original files from the workspace (issue #90). Every imported
document's ⋯ menu now offers Export original file, which saves the stored
original — PDF, Word, recording, photo, any format — byte-for-byte to a
location you choose, even when the file you imported it from is long gone.
Imported documents previously had no export action at all (only generated
documents did). The export shows the encryption-boundary warning first (the
saved copy is not protected by your workspace password), writes atomically,
and records only the document id in the activity log. One document per export
for now; bulk export is a possible follow-up.
Changed
- Qwen3.8 27B is the new recommended chat model for 24 GB and ≥32 GB machines
(owner ratification 2026-08-16;docs/model-benchmarks.md§9.4). Three new
catalog entries (Q4_K_M, Q5_K_M, and the selectable Q6_K quality ceiling for
24 GB GPUs), all license-reviewed with verified upstream hashes. The Qwen3.6
27B pair stays in the catalog, still ranked and selectable. - Every screen shares one content width (issues #166/#171). All screens —
Home, Documents, Translate, Images, AI model, Skills, Settings — now use the
same centred content width with symmetric side gutters, so pages no longer
read as differently sized; only the chat keeps its full-bleed layout with the
centred transcript. Explanatory texts inside are capped at a readable line
length, so the wider pages don't make help texts harder to read. The centred
content also no longer shifts sideways by the scrollbar's width when
switching...
v0.1.56
Which file do I need?
The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:
| Your computer | Download this ONE file |
|---|---|
| Windows 10/11 | HilbertRaum-0.1.56-portable.exe |
| macOS (Apple Silicon) | HilbertRaum-0.1.56-mac-arm64.app.zip |
| Linux | HilbertRaum-0.1.56.AppImage |
The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.
This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.
Added
- Local chat — a
llama.cppruntime running GGUF models entirely on-device
(CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
Granite) and an on-machine benchmark that recommends the best-fit model for
available RAM. A built-in demo mode runs the whole UI with no model files
and no network. - Document Q&A with citations — import PDF / Word / text, ask questions, and
get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
reranker, scoped by collections. - Image understanding — ask questions about a picture with a local vision
model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
and deletable — per-entry and total sizes are shown, a confirmed "Clear
history" removes everything at once (issue #122), and the history's disk
footprint appears in Settings → Diagnostics. - Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
on-device OCR ("Make searchable") on scanned pages (bundled German + English
language files; no cloud OCR). - Document tasks & skills — summarize, translate, and compare documents;
install reusable skills for structured extraction (bank statements,
invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe). - Evidence packs (review mode) — review a document-grounded answer block by
block against its sources, record explicit decisions and notes, and export the
review as a self-contained evidence pack in HTML or PDF — generated
locally and offline, with honest coverage, freshness, and limitation notes. A
pack supports human review; it is not a correctness certification. - Translate view + dedicated translation model (TranslateGemma) — a top-level
Translate screen for live text translation and drag-and-drop document
translation across 51 languages, source and target (the model's full
production tier — from German and English to Arabic, Chinese, Swahili, and
Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
sidecar (downloaded on demand behind the license acknowledgement; not
bundled) — never the chat model; document translations materialize as
searchable, exportable local documents. GPU-accelerated when the machine
allows it, with an automatic CPU fallback. Calibrated against the real model
(per-language round-trip evidence + measured tokenizer weights) so a window
can only over-chunk, never overflow. - Encrypted, portable workspace — an optional password-encrypted workspace
(AES-256-GCM with Argon2id key derivation) covering the database, imported-document
copies, and the diagnostics log; keep models plus the workspace on an external
drive and move between laptops. - Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
context-isolated renderer; a strict Content-Security-Policy; deny-by-default
renderer permissions; an offline guard that trips on any non-loopback connection
attempt; and a tamper-evident local audit log (ids/counts only, never content). - Cross-platform & distribution — Windows-first, with macOS and Linux supported
in the architecture; portable / preconfigured-drive distribution via
scripts/build-commercial-drive.*. - Standard project docs — this
CHANGELOG.mdand a
CODE_OF_CONDUCT.md. - Windows engine build attached to releases (issue #102 item 1). Each release
now carriesllama-runtime-win-x64.zipbeside the existing mac Metal zip: the
same SHA-256-verified llama.cpp build the in-app installer fetches, packaged so
it unzips straight into the drive'sruntime/llama.cpp/folder — an offline /
air-gapped install path that needs no repo scripts. - Opt-in perf mark log for measurement runs: with
HILBERTRAUM_PERF_LOG=1set in
the environment, the app appends timestamped timing marks (startup phases, vault
unlock/lock split, model checksum vs. load split, time to first token, document
ingestion phases) tologs/perf.log. Off by default; no file is created without the
variable. Records the cold-load figuresdocs/model-benchmarks.md§11.4 lists as
still missing. Seedocs/benchmark.md"Perf marks". - Model checksum passes are visible now (issue #106). Verifying a model's
multi-GB file — which can take minutes from a slow USB stick and used to run
with no trace anywhere — now writes one diagnostics-log line per real
verification (which model, how many bytes, how long), pluschecksum_start/
checksum_donemarks in the opt-in perf log. Overlapping verifications of the
same file (for example a start racing an AI Model screen visit) also no longer
each read the whole file — they share one pass. - Honest progress while a model starts (issue #107). Once the app has
measured your drive's real read speed, the Chat screen's "your model is
starting" panel shows the model file size and an approximate percentage read
instead of an indefinite spinner. Fresh installs (no measurement yet) keep the
plain message rather than showing a made-up number. - Cold model starts are much faster on slow drives (issue #114). While a
model loads, the app now also reads the model file front-to-back in the
background, which lines the operating system's file cache up with what the
load is about to need. Measured on a 16 GB laptop: a 6.65 GB model from a slow
USB stick started in 5:46 instead of 11:20 (−49%), and from a portable SSD in
10 s instead of 15 s (−36%). Starts that are already fast stay as they are —
the helper is skipped when the file was just verified (its content is already
in the cache) and stops the moment the load finishes or is cancelled. - One-click skill offers on document answers (issue #80). When a document
answer cannot serve the shape you asked for — "categorize the transactions
and sum per category" answered by an engine that can only list values — the
answer now carries a one-click "Run "Bank Statement Analysis" for this
question" action instead of only a prose pointer; clicking re-answers the
same question with the skill (nothing ever runs without the click). On the
residue the deterministic router provably cannot classify, a bounded,
grammar-constrained on-device model pick (marked "Suggested by the local
model") may fill the same offer; ordinary questions keep their zero-model-call
routing, and every failure quietly falls back to today's answer. - Export your original files from the workspace (issue #90). Every imported
document's ⋯ menu now offers Export original file, which saves the stored
original — PDF, Word, recording, photo, any format — byte-for-byte to a
location you choose, even when the file you imported it from is long gone.
Imported documents previously had no export action at all (only generated
documents did). The export shows the encryption-boundary warning first (the
saved copy is not protected by your workspace password), writes atomically,
and records only the document id in the activity log. One document per export
for now; bulk export is a possible follow-up.
Changed
- Every screen shares one content width (issues #166/#171). All screens —
Home, Documents, Translate, Images, AI model, Skills, Settings — now use the
same centred content width with symmetric side gutters, so pages no longer
read as differently sized; only the chat keeps its full-bleed layout with the
centred transcript. Explanatory texts inside are capped at a readable line
length, so the wider pages don't make help texts harder to read. The centred
content also no longer shifts sideways by the scrollbar's width when
switching between scrolling and non-scrolling pages. - Document translations pack fuller windows on compound-heavy languages
(issue #165). The window planner now budgets and fills in the same unit
(words), so German/Czech-style prose no longer splits into 1.5–2.5× more
windows than needed — fewer model calls and materially less wall-clock on
exactly the languages whose windows decode sl...
v0.1.55
This download is the app only — the AI models are fetched separately (README step 2 / the in-app downloader). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.
Added
- Local chat — a
llama.cppruntime running GGUF models entirely on-device
(CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
Granite) and an on-machine benchmark that recommends the best-fit model for
available RAM. A built-in demo mode runs the whole UI with no model files
and no network. - Document Q&A with citations — import PDF / Word / text, ask questions, and
get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
reranker, scoped by collections. - Image understanding — ask questions about a picture with a local vision
model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
and deletable. - Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
on-device OCR ("Make searchable") on scanned pages (bundled German + English
language files; no cloud OCR). - Document tasks & skills — summarize, translate, and compare documents;
install reusable skills for structured extraction (bank statements,
invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe). - Evidence packs (review mode) — review a document-grounded answer block by
block against its sources, record explicit decisions and notes, and export the
review as a self-contained evidence pack in HTML or PDF — generated
locally and offline, with honest coverage, freshness, and limitation notes. A
pack supports human review; it is not a correctness certification. - Translate view + dedicated translation model (TranslateGemma) — a top-level
Translate screen for live text translation and drag-and-drop document
translation across 51 languages, source and target (the model's full
production tier — from German and English to Arabic, Chinese, Swahili, and
Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
sidecar (downloaded on demand behind the license acknowledgement; not
bundled) — never the chat model; document translations materialize as
searchable, exportable local documents. GPU-accelerated when the machine
allows it, with an automatic CPU fallback. Calibrated against the real model
(per-language round-trip evidence + measured tokenizer weights) so a window
can only over-chunk, never overflow. - Encrypted, portable workspace — an optional password-encrypted workspace
(AES-256-GCM with Argon2id key derivation) covering the database, imported-document
copies, and the diagnostics log; keep models plus the workspace on an external
drive and move between laptops. - Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
context-isolated renderer; a strict Content-Security-Policy; deny-by-default
renderer permissions; an offline guard that trips on any non-loopback connection
attempt; and a tamper-evident local audit log (ids/counts only, never content). - Cross-platform & distribution — Windows-first, with macOS and Linux supported
in the architecture; portable / preconfigured-drive distribution via
scripts/build-commercial-drive.*. - Standard project docs — this
CHANGELOG.mdand a
CODE_OF_CONDUCT.md.
Changed
- Deep-index extraction is more reliable under reasoning-prone models — the
"Build deep index" structured-extract pass now grammar-constrains the model's
reply (the same JSON-schema mechanism the bank-statement categorizer uses), so
sections can no longer come back unreadable because the model answered in
prose or code fences; "unparsed" sections in listing answers should now be
rare, and the existing retry/salvage safety net is unchanged (wave STR-1; see
thedocs/architecture.md"Skills & tools architecture review (2026-07-19) —
design record").
Fixed
- Scanned-PDF OCR is startable again from the Documents row (it had become
unreachable after a row-actions refactor): "Make searchable (OCR)" is an inline
button on the scan's row, already-recognized PDFs can be re-run via
"Read again (OCR)", Translate now explains scanned PDFs (make searchable
first) instead of calling them unsupported, and progress is honest through the
final "Finishing" step. Packaged builds no longer carry the dev-only localhost
CSP relaxation in their HTML meta tags (wave OCR-R, PR #75; see the
docs/architecture.md"OCR audit (2026-07-18) — remediation ledger"). npm run devno longer 500s on the first page load — a false "no CSP meta
tag" throw during dev serve (the guard mis-read a deliberate, byte-identical
no-op rewrite as a missing tag) is fixed; packaged builds were never affected
(wave DEP-1, PR #77).
Security
- Post-DEP-1 advisory batch cleared (wave DEP-2, 2026-07-23) — five transitive
dev/build-tooling packages patched lockfile-only, semver-compatible, no manifest
changes: node-tar 7.5.21 (decompression/parse DoS CRITICAL + infinite-loop,
PAX-crash and NUL-byte advisories), js-yaml 4.3.0 (merge-key quadratic CPU;
electron-builder's copy — the app's own manifest parser is the separateyaml
package and was never affected), fast-uri 3.1.4 (host confusion ×2),
brace-expansion 1.1.16 / 2.1.2 / 5.0.8 (exponential-time{}expansion DoS,
CVE-2026-13149 — Dependabot alert #56), and DOMPurify 3.4.12 (the one
production-scope member: streamdown → mermaid ships it in the renderer;
CUSTOM_ELEMENT_HANDLINGsanitizer bypass, low severity, config not used by
mermaid).npm audit: 0 vulnerabilities again. - All critical- and high-severity Dependabot alerts cleared (wave DEP-1, PR #77)
— Vitest 3.2.6 (CVE-2026-47429, UI-server arbitrary file read/execute), Electron
39.8.10 (command-line switch injection, four use-after-free classes,
permission-origin confusion, header injection), Vite 6.4.3 + electron-vite 3.1.0
(server.fs.denybypass on Windows, path traversal), form-data 4.0.6 (CRLF
injection, CVE-2026-12143), undici 7.28.0 + 6.27.0 (TLS-bypass and cross-origin
routing via a SOCKS5 proxy, header injection), and esbuild 0.25.12 (dev-server
CORS).npm auditnow reports 0 vulnerabilities. Packaged-build security was
re-verified on the new Electron 39 runtime: the strict Content-Security-Policy
response header still attaches and enforces onfile://in both windows (see
docs/architecture.md"Dependency remediation — design record (wave DEP-1,
PR #77)").
v0.1.54
This download is the app only — the AI models are fetched separately (README step 2 / the in-app downloader). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.
Added
- Local chat — a
llama.cppruntime running GGUF models entirely on-device
(CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
Granite) and an on-machine benchmark that recommends the best-fit model for
available RAM. A built-in demo mode runs the whole UI with no model files
and no network. - Document Q&A with citations — import PDF / Word / text, ask questions, and
get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
reranker, scoped by collections. - Image understanding — ask questions about a picture with a local vision
model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
and deletable. - Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
on-device OCR ("Make searchable") on scanned pages (bundled German + English
language files; no cloud OCR). - Document tasks & skills — summarize, translate, and compare documents;
install reusable skills for structured extraction (bank statements,
invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe). - Evidence packs (review mode) — review a document-grounded answer block by
block against its sources, record explicit decisions and notes, and export the
review as a self-contained evidence pack in HTML or PDF — generated
locally and offline, with honest coverage, freshness, and limitation notes. A
pack supports human review; it is not a correctness certification. - Translate view + dedicated translation model (TranslateGemma) — a top-level
Translate screen for live text translation and drag-and-drop document
translation across 51 languages, source and target (the model's full
production tier — from German and English to Arabic, Chinese, Swahili, and
Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
sidecar (downloaded on demand behind the license acknowledgement; not
bundled) — never the chat model; document translations materialize as
searchable, exportable local documents. GPU-accelerated when the machine
allows it, with an automatic CPU fallback. Calibrated against the real model
(per-language round-trip evidence + measured tokenizer weights) so a window
can only over-chunk, never overflow. - Encrypted, portable workspace — an optional password-encrypted workspace
(AES-256-GCM with Argon2id key derivation) covering the database, imported-document
copies, and the diagnostics log; keep models plus the workspace on an external
drive and move between laptops. - Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
context-isolated renderer; a strict Content-Security-Policy; deny-by-default
renderer permissions; an offline guard that trips on any non-loopback connection
attempt; and a tamper-evident local audit log (ids/counts only, never content). - Cross-platform & distribution — Windows-first, with macOS and Linux supported
in the architecture; portable / preconfigured-drive distribution via
scripts/build-commercial-drive.*. - Standard project docs — this
CHANGELOG.mdand a
CODE_OF_CONDUCT.md.
Fixed
- Scanned-PDF OCR is startable again from the Documents row (it had become
unreachable after a row-actions refactor): "Make searchable (OCR)" is an inline
button on the scan's row, already-recognized PDFs can be re-run via
"Read again (OCR)", Translate now explains scanned PDFs (make searchable
first) instead of calling them unsupported, and progress is honest through the
final "Finishing" step. Packaged builds no longer carry the dev-only localhost
CSP relaxation in their HTML meta tags (wave OCR-R, PR #75; see the
docs/architecture.md"OCR audit (2026-07-18) — remediation ledger"). npm run devno longer 500s on the first page load — a false "no CSP meta
tag" throw during dev serve (the guard mis-read a deliberate, byte-identical
no-op rewrite as a missing tag) is fixed; packaged builds were never affected
(wave DEP-1, PR #77).
Security
- All critical- and high-severity Dependabot alerts cleared (wave DEP-1, PR #77)
— Vitest 3.2.6 (CVE-2026-47429, UI-server arbitrary file read/execute), Electron
39.8.10 (command-line switch injection, four use-after-free classes,
permission-origin confusion, header injection), Vite 6.4.3 + electron-vite 3.1.0
(server.fs.denybypass on Windows, path traversal), form-data 4.0.6 (CRLF
injection, CVE-2026-12143), undici 7.28.0 + 6.27.0 (TLS-bypass and cross-origin
routing via a SOCKS5 proxy, header injection), and esbuild 0.25.12 (dev-server
CORS).npm auditnow reports 0 vulnerabilities. Packaged-build security was
re-verified on the new Electron 39 runtime: the strict Content-Security-Policy
response header still attaches and enforces onfile://in both windows (see
docs/architecture.md"Dependency remediation — design record (wave DEP-1,
PR #77)").
v0.1.53
This download is the app only — the AI models are fetched separately (README step 2 / the in-app downloader). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.
The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.
Added
- Local chat — a
llama.cppruntime running GGUF models entirely on-device
(CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
Granite) and an on-machine benchmark that recommends the best-fit model for
available RAM. A built-in demo mode runs the whole UI with no model files
and no network. - Document Q&A with citations — import PDF / Word / text, ask questions, and
get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
reranker, scoped by collections. - Image understanding — ask questions about a picture with a local vision
model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
and deletable. - Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
on-device OCR ("Make searchable") on scanned pages (bundled German + English
language files; no cloud OCR). - Document tasks & skills — summarize, translate, and compare documents;
install reusable skills for structured extraction (bank statements,
invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe). - Evidence packs (review mode) — review a document-grounded answer block by
block against its sources, record explicit decisions and notes, and export the
review as a self-contained evidence pack in HTML or PDF — generated
locally and offline, with honest coverage, freshness, and limitation notes. A
pack supports human review; it is not a correctness certification. - Translate view + dedicated translation model (TranslateGemma) — a top-level
Translate screen for live text translation and drag-and-drop document
translation across 51 languages, source and target (the model's full
production tier — from German and English to Arabic, Chinese, Swahili, and
Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
sidecar (downloaded on demand behind the license acknowledgement; not
bundled) — never the chat model; document translations materialize as
searchable, exportable local documents. GPU-accelerated when the machine
allows it, with an automatic CPU fallback. Calibrated against the real model
(per-language round-trip evidence + measured tokenizer weights) so a window
can only over-chunk, never overflow. - Encrypted, portable workspace — an optional password-encrypted workspace
(AES-256-GCM with Argon2id key derivation) covering the database, imported-document
copies, and the diagnostics log; keep models plus the workspace on an external
drive and move between laptops. - Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
context-isolated renderer; a strict Content-Security-Policy; deny-by-default
renderer permissions; an offline guard that trips on any non-loopback connection
attempt; and a tamper-evident local audit log (ids/counts only, never content). - Cross-platform & distribution — Windows-first, with macOS and Linux supported
in the architecture; portable / preconfigured-drive distribution via
scripts/build-commercial-drive.*. - Standard project docs — this
CHANGELOG.mdand a
CODE_OF_CONDUCT.md.