Skip to content

Releases: HilbertraumAI/HilbertRaum

v0.1.62

Choose a tag to compare

@github-actions github-actions released this 27 Sep 23:51
143f332

Which file do I need?

The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:

Your computer Download this ONE file
Windows 10/11 HilbertRaum-0.1.62-portable.exe
macOS (Apple Silicon) HilbertRaum-0.1.62-mac-arm64.app.zip
Linux HilbertRaum-0.1.62.AppImage

The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.

This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Two optional downloads are offered in the app only when you need them: the knowledge-pack tools, and the text-recognition (OCR) files for scanned PDFs and photos (German + English, about 4 MB; offered on the scan or photo in Documents and on the AI Model screen). Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

Added

  • Text recognition (OCR) can now be added from inside the app. When a scanned PDF or a
    photo of a page needs the OCR language files and they are not on the drive, its row in
    Documents offers Download OCR files — and the AI Model screen shows the same
    offer as "Text recognition for scans and photos (optional)". A short confirmation names the
    two files (German and English, about 4 MB, Apache-2.0) and where they come from; once they
    are downloaded and checked, text recognition works without restarting the app: the scan
    offers Make searchable (OCR), and a photo that failed reads with Try again. Like
    every other download it asks first, needs the drive policy and Allow internet access… to
    permit it, and is verified before use. Portable-app users no longer need the drive-setup
    scripts for OCR; they remain the offline alternative (#410).
  • While you dictate, the app tells you at once when nothing is reaching the microphone.
    After about two seconds without any sound, a note appears under the message box and
    disappears as soon as sound comes in — so a muted or wrong microphone is caught before you
    stop, not after (#497).

Changed

  • The AI engine is updated to llama.cpp b11146 (the upstream v0.5.0 release). In our
    document-question tests the ten built-in chat models answer as accurately as before. On the
    Qwen3.5, Qwen3.6, Qwen3.8 and Gemma 4 models, coming back to a conversation after something else used the model (a
    document being indexed, another conversation) no longer means re-reading the whole
    conversation first: the engine now restores it from memory, so that first reply starts sooner.
    The memory the engine keeps for this is limited to an eighth of the computer's RAM, at most
    8 GB; when it is full, the oldest saved conversation is dropped. A drive set up with the
    previous engine keeps using it until the drive-setup script is run again (#512).
  • Document translation works in slightly smaller parts. A long document is now translated
    in parts of about 640 words instead of about 690, so a part full of dense text, such as an
    invoice or a table, stays within the input size the translation model was trained on. A long
    document takes about 7 % more parts (#512).
  • Dictation now ignores background noise. Before transcribing what you dictated, the app
    runs a small voice-activity model (Silero VAD) so a recording with no speech in it yields
    "No speech was recognized" instead of a stray word. The model is a second, small file of the
    speech model: the AI Model screen fetches it together with the speech model, and a drive
    that has only the older single file shows the speech model as incomplete until that one
    file is added (#504).
  • The microphone button no longer disappears when the speech model is missing. On a
    drive without the speech model, the chat composer now shows the button greyed out; clicking
    it explains what is missing and opens the AI Model screen, where the model can be added
    (#497).

Fixed

  • A model that cannot be loaded is now reliably told apart from a graphics-card problem.
    Before, on some systems such a model could switch graphics acceleration off for every model,
    or two different start failures could be taken for the same one (#515).
  • Inline math written as $…$ now renders in chat. Before, only $$…$$ blocks and
    \(…\) spans were typeset, so the inline formulas many local models write showed up as raw
    text with the dollar signs. A single-dollar span is typeset only when it looks like math, so
    amounts such as "$5 and $10" still read as plain text (#501).
  • The "reply cut off" note no longer blames the context window for every cut. A grounded
    answer from your documents or a knowledge pack is stopped by a fixed length the app allows
    for one reply, and the note used to say the model's context limit was reached and to suggest
    raising the context size, which cannot help there. The note now just says the reply was cut
    off, and its tooltip suggests raising the context size only when the context really ran out
    (#498).
  • Superscripts and subscripts in knowledge-pack articles stay readable everywhere. An area
    of 25 m² or a number like 10⁶ in an ordinary paragraph used to reach the model and the
    article view fused as 25 m2 and 106; only table text kept them apart. Paragraphs and
    headings now use the same readable form as tables (25 m^2, 10^6, H_2O), and search
    still finds them from a plainly typed question such as "CO2" or "106" (#488).
  • Voice dictation works right after you install the speech model or the voice engine — no
    restart.
    The app used to keep its startup answer ("not available") until it was restarted,
    without saying so; the microphone button now appears as soon as the download or install
    finishes, also when you stayed in the chat meanwhile (#497).
  • A silent recording no longer turns into a stray word. When nothing reaches the
    microphone — it is muted, the wrong input device is selected, or the operating system does
    not allow the app to use it — dictation now says so instead of inserting the word "you"
    (#497).
  • Two small article-reading fixes: a rare text leak, and photo captions. Reading a
    Wikipedia table that contained a certain very rare block of raw text could leak a stray
    fragment of it into what the app reads back; that block is now skipped correctly, the same
    way it already is outside of tables. And a photo's caption, which used to be dropped along
    with the photo itself, is now kept as part of the article's text — the app reads captions on
    purpose now, the way it already reads a table's own caption.

Issues resolved

  • #410 — OCR language files are the only drive asset with no in-app install path (follow-up to #59)
  • #436 — Accessibility: ErrorBanner's failure message is never announced — the nested inner role="status" defeats the always-mounted alert (M-U1)
  • #460 — Tests: 99 of 129 openDatabase() suites never close the handle — ~1,450 locked temp roots per Windows run
  • #488 — Superscripts and subscripts are readable in table text but still fuse in prose
  • #493 — A CDATA section inside a kept table can still leak its text, where ordinary paragraphs do not
  • #498 — Cut-off badge and its remedy hint blame the context window for the app's own cap
  • #501 — Chat: inline math written as $...$ shows as raw text, only $$...$$ blocks render
  • #504 — Silero VAD for whisper: fetched in-app with the speech model, used for dictation (follow-up to #497)
  • #512 — Bump the llama.cpp runtime pin b9849 → b11146 (v0.5.0)
  • #515 — Start ladder: the #312 model-vs-device check compares non-comparable log lines (colour reset / per-process timestamps)
  • #517 — Grounded-QA scorer misses German abstentions ("nennt … nicht", "nicht festgelegt", "nicht direkt genannt")
  • #518 — Re-capture the --list-devices and vision SSE fixtures on the current pin (F-40)

What's Changed

  • fix(zim): CDATA in kept tables, and figure captions (#493) by @comilionas in #500
  • fix(dictation): restart-free speech-model activation, silence gate before whisper, visible "not installed" mic (#497) by @comilionas in #503
  • feat(chat): live "no sound is reaching the microphone" hint while dictating (#497 follow-up) by @comilionas in #505
  • feat(dictation): Silero VAD before whisper for dictation, fetched in-app with the speech model (#504) by @comilionas in #506
  • fix: readable sup/sub in knowledge-pack prose (#488), cause-neutral cut-off badge (#498), inline $…$ math (#501) by @comilionas in #509
  • OCR language files install in-app and activate without a restart (#410) by @comilionas in #510
  • Translation planner constants 2.8 in / 3.1 out (#512 decision 3) by @comilionas in #513
  • llama.cpp b...
Read more

v0.1.61

Choose a tag to compare

@github-actions github-actions released this 21 Sep 11:10
7169f9e

Which file do I need?

The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:

Your computer Download this ONE file
Windows 10/11 HilbertRaum-0.1.61-portable.exe
macOS (Apple Silicon) HilbertRaum-0.1.61-mac-arm64.app.zip
Linux HilbertRaum-0.1.61.AppImage

The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.

This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

Added

  • Knowledge-pack answers can now draw on the values inside a page's tables. Material property
    tables, comparison tables and similar structured content used to be dropped entirely when a
    knowledge-pack page was read; those values now reach the pack's content and can be cited in an
    answer, the same way ordinary paragraph text already was.
  • A select-all checkbox in the documents view. Next to the per-row checkboxes and the "Add to
    project" action, one control now selects or clears every document in the current list at once
    (#214).

Changed

  • Grounded document and knowledge-pack answers no longer vary in wording between two askings of
    the same question.
    The app now sends the model a fixed decoding setting (no randomness, and a
    per-pass length cap that is continued rather than cut off) instead of the model server's own
    defaults, so a repeated question gets the same wording back. A related case, where the cut-off
    notice itself can be inaccurate, is tracked separately (#498).
  • Knowledge-pack answers only name a pack's language when it is relevant. The search planner
    used to mention the pack's archive language on every question; it now does so only when the
    question is asked in a different language than the pack itself. On a maintainer benchmark this
    roughly doubled how often an English question against a German-language pack reached the right
    article (27 in 100 to 50 in 100), with German-language questions unaffected (#486).

Fixed

  • Knowledge-pack table content no longer leaks a page's stylesheet or repeats a formula. A
    style/script element or a mathematical formula nested inside a table that a knowledge pack reads
    could end up in the delivered text — raw CSS in one case, both the rendered and the source form
    of the same formula in the other. Both are now handled the way the rest of the page already was:
    styling is dropped, and a formula is converted once, not twice (#485, #490).
  • The reranker recovers from a graphics-card failure instead of staying off for the session.
    When ranking on the graphics card fails, the app now falls back to the processor for the rest of
    the session, and the request-size ceiling that fallback enforces is correctly re-applied along
    every path that can trigger it (#474).
  • The document-search model's shutdown can no longer overlap with itself. (#475)
  • The Performance page now accounts for a loaded ranking model. Its graphics-memory figures
    used to check only the chat and translation models; they now include a resident ranking model
    and the graphics memory it is contending for (#476).
  • The counter behind a pending model switch can no longer silently read as empty. (#477)
  • A test covering evidence-pack PDF text extraction no longer misreports a passing result as a
    failure.
    (#481)
  • The release build now runs its test suite split into shards, the same way ordinary CI already
    does, instead of one long unsharded run.
    (#458)

Issues resolved

  • #214 — Add a "Select all" checkbox in document upload view
  • #447 — The ZIM query expander was never measured against the #399 prompt-cache eviction (inference, not measurement)
  • #458 — CI: both Windows legs run at their time budget and flake — three re-run cycles in two days
  • #474 — Reranker recovery after a GPU failure, including a GPU→CPU demotion
  • #475 — The E5 embedder's teardown overlap across workspace lock and quit (pre-existing)
  • #476 — The Performance card summary line omits a GPU-resident reranker
  • #477 — Make the reranker's pending-model-switch counter impossible to leave unwired
  • #478 — Deliver tables to the model instead of dropping them
  • #481 — evidence-pack-pdf-smoke: the two archive-pack cases fail on master where the font splits the "fi" ligature (skipped in CI)
  • #485 — Stylesheet source can leak into delivered table text and into citation snippets
  • #486 — English questions retrieve far less than German ones from a German Wikipedia pack
  • #490 — Formulas inside tables are delivered twice, as loose markup characters plus raw TeX source

What's Changed

  • eval(#447): the ZIM query expander evicts the prefix on every pack-scoped turn — measured by @humaniser in #480
  • feat(zim): deliver tables to the model instead of dropping them by @comilionas in #479
  • docs: committed eval evidence under eval/results/ is not "logs" in the hard rule by @comilionas in #484
  • docs: record the knowledge-pack retrieval research outcome and its negative result by @comilionas in #489
  • fix(zim-tables): stop a style/script body and a doubled math formula from leaking into kept table text by @comilionas in #492
  • fix(rag): pin grounded answers to a fixed decoding setting by @comilionas in #491
  • UI select-all, PDF-smoke ligature fix, and release.yml Windows sharding by @comilionas in #494
  • Reranker recovery, embedder teardown, Performance card, and a required-wiring fix by @comilionas in #495
  • zim: plan in a pack's language only when the question's differs (#486) by @comilionas in #496
  • docs: release notes, shipping rule and BUILD_STATE drain before the next tag by @comilionas in #499
  • release: v0.1.61 — version bump by @comilionas in #502

Full Changelog: v0.1.60...v0.1.61

v0.1.60

Choose a tag to compare

@github-actions github-actions released this 18 Sep 18:17
5a4894a

Which file do I need?

The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:

Your computer Download this ONE file
Windows 10/11 HilbertRaum-0.1.60-portable.exe
macOS (Apple Silicon) HilbertRaum-0.1.60-mac-arm64.app.zip
Linux HilbertRaum-0.1.60.AppImage

The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.

This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

Added

  • Knowledge packs, one click from the Documents title. The Documents screen now has a switch
    next to its title — My documents | Knowledge packs — instead of a "Reference" entry at the
    bottom of the sidebar. In the packs view each pack has an Enabled switch, an Ask this
    pack
    button that opens a chat answering from that pack alone, and a ⋯ menu for Remove.
    The first-run notice became a setup card, and the empty view names library.kiwix.org with a
    Copy the library address button. Home's readiness card gained a Knowledge packs row,
    and the chat's "Answering from" picker offers Add packs… while you have none.
  • Filter documents by name. A search box above the document list narrows the current
    section as you type.
  • Check all model files, and stop it when you want to. A new action near the top of the AI
    Model screen checks every model file on the drive against its published checksum. It replaces a
    check the app used to run on its own every time you opened that screen (see Changed). On a slow
    drive the full check can take several minutes, so Stop checking sits beside its progress bar:
    the files it already checked stay checked, the rest are simply left unchecked, and stopping is
    not treated as a failure. The check keeps running if you move to another screen — come back and
    the progress bar and Stop checking are still there.
  • A quieter left navigation. The sidebar now has three groups: Chat, Documents, Translate
    and Images for everyday work; AI Model and Performance for the machine; Settings at the
    bottom. The HilbertRaum mark at the top is the Home button and lights up when you are on
    Home, so the separate Home entry is gone. Skills moved into Settings as its own tab; the
    skill picker in the chat composer is unchanged.
  • A new Performance page in the left navigation. It answers, in plain words, what this
    computer can run and how fast: one sentence, then four tiles for speed, memory, graphics
    memory and drive, each with a rating word, and a "Your model" line that says where your
    current model lands on this computer: in graphics memory, partly on the graphics card with
    the rest in RAM, on the processor, or too large, with the sizes behind the verdict once the
    model has started. Apple Silicon shows one unified memory figure. A "Models on this computer"
    card lists every model the app can hold (chat, translation, images, document search, voice),
    where each runs and whether it is loaded right now, and sums what the graphics card and the
    RAM would need with everything loaded at once. Below it, figures from real use (your last answer, the last model
    start, the last file check) and one result for every computer this drive has been plugged
    into. The check runs on its own the first time and whenever the drive lands on a different
    computer; on a computer it already knows, the earlier result is restored so the recommended
    model follows the machine. While a check runs you see its steps instead of a plain "Running…"
    button, and if no model has run yet the page offers to start the recommended one and measure.
    The technical table stays on Settings → Diagnostics. English and German.
  • A new computer's first check never borrows another computer's drive-speed reading.
    Plugging this drive into a computer it hasn't seen before measures that computer's own read
    speed from scratch; a warning about a slow drive, or the lack of one, always reflects this
    computer, never whichever computer the drive was in last.
  • The Performance page refreshes itself and stays precise about what it measured. It updates
    in place the moment a check finishes, a model starts, or a file is verified, with no need to
    leave and come back. A speed reading counted from streamed chunks rather than the model's own
    timing is marked Approximate everywhere it appears, including another computer's row and the
    copied report, which also always names the computer it describes. The graphics-memory tile
    says plainly when a chip's memory is Integrated and shared with the rest of the computer, and
    shows Not recorded, instead of guessing, for a computer this drive visited before the check
    could record its card. When a model only partly fits in graphics memory, the explanation now
    states the exact safety margin the AI engine reserves. After a check you started with the
    keyboard, focus returns to Check again.
  • Checking your computer's speed no longer competes with starting a model. The automatic
    first-time check now waits for a model that is already starting up to finish loading before it
    measures your drive, instead of reading and loading at the same time.
  • Knowledge packs: ask an offline Wikipedia (or any ZIM archive). Register ZIM
    files — e.g. from the Kiwix library — as knowledge packs (Documents → Knowledge
    packs, or just drop them into the drive’s zim/ folder), tick them as sources in a
    documents chat, and answers draw on them with citations that name the archive and
    open the article offline. Fully local: the pack server binds to 127.0.0.1 only,
    archives are used in place and never copied. Needs the kiwix-tools binaries on the
    drive — still a manual step in this release (see the user guide §7b).
  • Keep a knowledge-pack article in your documents. Save to my documents files a copy of a
    cited article alongside everything else you have imported, so it stays searchable even after
    you disable or remove that pack, or unplug the drive it lives on. The action sits on the
    citation itself, next to Open article, and inside the article view; it names the copy it
    filed, and saving the same article twice tells you it is already there instead of making a
    second one. Reviews stay read-only, so it is not offered there.
  • Knowledge packs are found once, not on every open. The list is discovered when you
    unlock and by an explicit Refresh in Documents → Knowledge packs, instead of a
    fresh drive scan every time the panel opens or a message is sent — packs load
    instantly, and copying a new archive onto the drive shows up after Refresh. A pack
    whose file was replaced by a different archive is now shown as such ("Different
    archive") instead of quietly answering from the wrong one. Locking or quitting the app
    stops the pack server and removes its small generated index file. Opening an alias article
    (a redirect entry — about half of a Wikipedia archive's titles) shows the article it points
    to instead of an error. A large article that the bundled Windows pack server occasionally
    delivers only in part is re-requested instead of being reported as unavailable or dropped from
    the answer. Known limit on Windows: an archive can only be added from a path
    without umlauts or accents (the bundled kiwix-manage cannot read such paths) — the drive's
    zim/ folder always works.
  • Evidence reviews now name the knowledge pack, not a same-named document. Reviewing an
    answer that cites a knowledge-pack article records the archive, the article and its pack
    id honestly, and shows identity as not verifiable against the workspace instead of
    matching it to a similarly named document — the review, its HTML/PDF evidence pack and
    the Markdown transcript export all name the pack. Reviews created on a pre-release
    knowledge-pack build that cite an archive must be re-run.
  • Adding several knowledge packs at once now reports exactly what happened. If some of
    the chosen archives could not be added, you are told how many were added and how many
    were not — never the technical reason a file failed. Answers and the article viewer now
    double-check with the pack server that it is still the same one that was running a moment
    ago, and retry once if it was restarted mid-question, instead of silently trusting
    whatever answers on that port. One privacy limit to know: other programs running under
    your own user account can read the enabled packs while the workspace is unlocked, through
    the pack server, which has no password of its own; locking or quitting the app stops it.
  • Answer from knowledge packs alone, and see what each one did. A new Search my
    documents
    toggle in the sources picker lets you answer from ticked knowledge packs
    only — files attached to the chat are still used either way — and every answer now
    shows a "Knowledge packs:" line ...
Read more

v0.1.59

Choose a tag to compare

@github-actions github-actions released this 21 Aug 14:58
b087ea6

Which file do I need?

The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:

Your computer Download this ONE file
Windows 10/11 HilbertRaum-0.1.59-portable.exe
macOS (Apple Silicon) HilbertRaum-0.1.59-mac-arm64.app.zip
Linux HilbertRaum-0.1.59.AppImage

The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.

This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

Fixed

  • A second running copy of the app can no longer destroy your encrypted workspace
    (issue #208).
    Starting the app while another copy was already running — the natural
    upgrade flow: launch the new version, then close the old one — could silently and
    permanently corrupt an encrypted workspace on Windows, and the damage only surfaced at
    the next unlock, looking like a wrong password. Three fixes ship together: the app now
    refuses to start a second copy (the running copy's window comes to the front instead);
    the workspace re-encryption on lock/quit refuses to overwrite the vault with anything
    that is not actually your database; and a workspace whose encrypted data is damaged now
    says so plainly at unlock — with backup guidance — instead of implying the password was
    wrong. See the new troubleshooting entry "Your password is correct, but the workspace
    data on the drive is damaged".
  • Failing to open the workspace database no longer leaves the decrypted file behind on the
    drive, and no longer holds an invisible open file handle to it.

Changed

  • Developer runs (npm run dev) now use their own app-data folder (…\@hilbertraum\ desktop-dev) instead of sharing the released app's production workspace. A workspace
    previously created from a dev run stays on disk under the old path; point the dev run at
    it with HILBERTRAUM_DRIVE_ROOT if you still need it.

Issues resolved

  • #208 — Encrypted workspace destroyed on disk: .enc now decrypts (with a valid tag) to random noise, and the unlock failure looks like a wrong password

What's Changed

Full Changelog: v0.1.58...v0.1.59

v0.1.58

Choose a tag to compare

@github-actions github-actions released this 20 Aug 20:19
9903c1d

Which file do I need?

The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:

Your computer Download this ONE file
Windows 10/11 HilbertRaum-0.1.58-portable.exe
macOS (Apple Silicon) HilbertRaum-0.1.58-mac-arm64.app.zip
Linux HilbertRaum-0.1.58.AppImage

The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.

This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

Added

  • Local API (optional, off by default). Other programs on the same computer can be
    allowed to use the running chat model over a loopback-only, OpenAI-compatible endpoint
    (http://127.0.0.1:4980/v1 — GET /v1/models and POST /v1/chat/completions, streaming
    or not, with JSON-schema-constrained output). It is off until switched on behind a consent
    dialog, requires an access key by default, never reaches the internet, exposes no documents
    or conversations, keeps no record of what was asked or answered, exists only while the
    workspace is unlocked, gives your own chat priority over any outside caller, and can be
    forbidden by drive policy. Tutorial, client examples and the wire contract:
    docs/local-api.md.
  • Faster answers from the large Qwen3.8 models on a capable GPU (issue #182). These model
    files always carried a small built-in "draft" head that the engine loaded and ignored; it is
    now used, worth a measured 38–45 % faster text generation on the reference machine. It
    engages only when one GPU has room for the model plus ~3.5 GiB, and any refusal falls back
    to exactly the previous behaviour — a model can be slower to start once, never broken.

Changed

  • New recommended chat models for 24 GB and ≥32 GB machines (issue #196). Unsloth removed
    the Qwen3.8 files the previous recommendations pinned, so those downloads had begun failing
    with a 404. Two things changed. A model whose upstream source has been withdrawn now says
    so on the AI Model screen
    , with the reason and the date, instead of offering a Download
    button that cannot work — and the drive-provisioning scripts skip it with a clear line
    instead of retrying a dead link. And the closest published successors were re-measured on
    the reference machine before taking over the tiers: Qwen3.8 27B UD-Q4_K_M at 24 GB and
    UD-Q5_K_M at ≥32 GB. Answer quality is unchanged within measurement noise; the 24 GB
    pick generates about 19 % slower than the withdrawn file, which is recorded in the
    catalog rather than glossed over. Already-downloaded copies of the old files keep
    working
    — they still verify, still start, and keep the speed-up above.

Fixed

  • Downloading "Gemma 4 12B Instruct QAT Q4" works again (issue #201). Google replaced that
    model file on their servers in July with a corrected version, at the same address. The app
    checks every download against the exact file it expects, so it was transferring the full ~7 GB
    and only then reporting a checksum failure — with nothing you could do about it. The catalog
    now points at the corrected file, which was measured against the old one first and answers
    identically. If you already downloaded this model, your copy will now report a checksum
    failure and needs downloading again
    (~7 GB) — press Download on the AI Model screen and it
    replaces itself. No other model in the catalog is affected: all the rest were re-checked
    against their sources in the same pass.
  • Deleting a document really deletes its stored copy after a drive changes letter
    (issue #188). The workspace recorded each imported copy by absolute path, so moving the
    drive between computers — or just getting a different drive letter — left every stored copy
    unreachable: "Delete document" removed the entry while the encrypted copy stayed on disk,
    and exporting or previewing the original failed. Paths are now resolved relative to the
    drive, and existing entries heal themselves the first time each document is opened.
  • Re-indexing or deleting a single document now tells you what happened (issue #194).
    The action worked but reported nothing at all — no confirmation, no progress, no error. It
    now shows a spinner on that row, confirms with the document's name, and names the document
    in the message if it fails. "Re-index all" already behaved this way.
  • One AI job at a time, across every part of the app (issues #185, #186). Starting the
    Diagnostics benchmark, or a skill run, could put a second job on the model while an answer
    was still streaming — each part of the app tracked "busy" separately and none of them
    agreed. They now share one signal. A speed measurement taken while something else was
    running is discarded rather than saved, so a contended reading can no longer push your
    machine's model recommendation down permanently; when that happens, the benchmark says so.

Security

  • Electron 39.8.10 → 43.4.0 (issue #179), clearing the last open dependency alert:
    GHSA-jmr9-qjv8-65gv (CVE-2026-56876, high) in extract-zip, which is abandoned at its last
    version and could only be cleared by moving Electron. Chromium 142 → 150, Node 22 → 24,
    SQLite 3.51.2 → 3.53.1. Minimum operating systems are unchanged (Windows 10+, macOS 12+)
    and no stored format changed, so existing workspaces open exactly as before. npm audit
    reports 0 vulnerabilities.

Issues resolved

  • #179 — Bump Electron 39 → 43 (removes the unpatchable extract-zip dependency, Dependabot #83)
  • #182 — Enable MTP speculative decoding for the Qwen3.8 chat manifests (+38-45% measured decode)
  • #185 — Benchmark has no re-entrancy or busy guard (surfaced by the local-API generation gate)
  • #186 — A skill run and a chat stream can generate concurrently (surfaced by the local-API generation gate)
  • #188 — Document export offers "Originaldatei exportieren" for documents it cannot export — and the stored copy may only look missing (absolute stored_path)
  • #194 — Single-document re-index gives no success or progress feedback, while "Re-index all" toasts and shows a progress bar
  • #196 — qwen3.8 manifests: upstream deleted the pinned GGUFs (Dynamic 3.0 restructure), all three download URLs return 404
  • #201 — gemma4-12b-it-qat-q4: upstream re-uploaded the GGUF (corrected vocabulary) — the pinned SHA-256 no longer matches, so downloads fail verification
  • #202 — Four chat manifests carry rounded, estimated download.size_bytes instead of the real byte count

What's Changed

  • docs(build-state): archive closed-wave entries (retention ritual) by @comilionas in #183
  • Local API endpoint wave (P1-P6) by @comilionas in #184
  • docs: §9.4 addendum — measured MTP findings for Qwen3.8 (not adopted) by @humaniser in #181
  • Electron 39 → 43: clear Dependabot #83 (unpatchable extract-zip) by @comilionas in #187
  • fix(ingestion): portable stored-copy resolver — "delete" now actually deletes (#188) by @comilionas in #189
  • fix(runtime): one shared occupancy span for the in-app model lanes (#185, #186) by @comilionas in #192
  • feat(runtime): MTP speculative decoding for the Qwen3.8 Q4/Q5 chat manifests (#182) by @comilionas in #191
  • fix(tests): the read-only stored-copy diagnostic — rebuilt, run, and its password path fixed (#190 phases 1–2) by @comilionas in #193
  • fix(docs-ui): single-document re-index/delete now report success, progress and a named failure (#194) by @comilionas in #195
  • docs(local-api): document the local API as a first-class release feature (O2 lifted) by @comilionas in #197
  • docs(build-state): record where the #196 fix actually landed (O6 citation correction) by @comilionas in #198
  • #196 successor wave: measured UD manifests, ranks ratified (generational handover restored) by @humaniser in #199
  • docs(changelog): per-version release notes — the file is the release page, so it has to be true by @comilionas in #200
  • docs: 2026-08-20 audit remediation — the #196 successor wave never reached the catalog prose by @comilionas in #203
  • docs(build-state): 1,968 -> 763 lines — the retention rule was draining the wrong section by @comilionas in #204
  • fix(catalog): re...
Read more

v0.1.57

Choose a tag to compare

@github-actions github-actions released this 17 Aug 00:28
bf26c07

Which file do I need?

The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:

Your computer Download this ONE file
Windows 10/11 HilbertRaum-0.1.57-portable.exe
macOS (Apple Silicon) HilbertRaum-0.1.57-mac-arm64.app.zip
Linux HilbertRaum-0.1.57.AppImage

The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.

This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.

Added

  • Local chat — a llama.cpp runtime running GGUF models entirely on-device
    (CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
    Granite) and an on-machine benchmark that recommends the best-fit model for
    available RAM. A built-in demo mode runs the whole UI with no model files
    and no network.
  • Document Q&A with citations — import PDF / Word / text, ask questions, and
    get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
    reranker, scoped by collections.
  • Image understanding — ask questions about a picture with a local vision
    model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
    and deletable — per-entry and total sizes are shown, a confirmed "Clear
    history" removes everything at once (issue #122), and the history's disk
    footprint appears in Settings → Diagnostics.
  • Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
    on-device OCR ("Make searchable") on scanned pages (bundled German + English
    language files; no cloud OCR).
  • Document tasks & skills — summarize, translate, and compare documents;
    install reusable skills for structured extraction (bank statements,
    invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe).
  • Evidence packs (review mode) — review a document-grounded answer block by
    block against its sources, record explicit decisions and notes, and export the
    review as a self-contained evidence pack in HTML or PDF — generated
    locally and offline, with honest coverage, freshness, and limitation notes. A
    pack supports human review; it is not a correctness certification.
  • Translate view + dedicated translation model (TranslateGemma) — a top-level
    Translate screen for live text translation and drag-and-drop document
    translation across 51 languages, source and target (the model's full
    production tier — from German and English to Arabic, Chinese, Swahili, and
    Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
    sidecar (downloaded on demand behind the license acknowledgement; not
    bundled) — never the chat model; document translations materialize as
    searchable, exportable local documents. GPU-accelerated when the machine
    allows it, with an automatic CPU fallback. Calibrated against the real model
    (per-language round-trip evidence + measured tokenizer weights) so a window
    can only over-chunk, never overflow.
  • Encrypted, portable workspace — an optional password-encrypted workspace
    (AES-256-GCM with Argon2id key derivation) covering the database, imported-document
    copies, and the diagnostics log; keep models plus the workspace on an external
    drive and move between laptops.
  • Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
    context-isolated renderer; a strict Content-Security-Policy; deny-by-default
    renderer permissions; an offline guard that trips on any non-loopback connection
    attempt; and a tamper-evident local audit log (ids/counts only, never content).
  • Cross-platform & distribution — Windows-first, with macOS and Linux supported
    in the architecture; portable / preconfigured-drive distribution via
    scripts/build-commercial-drive.*.
  • Standard project docs — this CHANGELOG.md and a
    CODE_OF_CONDUCT.md.
  • Windows engine build attached to releases (issue #102 item 1). Each release
    now carries llama-runtime-win-x64.zip beside the existing mac Metal zip: the
    same SHA-256-verified llama.cpp build the in-app installer fetches, packaged so
    it unzips straight into the drive's runtime/llama.cpp/ folder — an offline /
    air-gapped install path that needs no repo scripts.
  • Opt-in perf mark log for measurement runs: with HILBERTRAUM_PERF_LOG=1 set in
    the environment, the app appends timestamped timing marks (startup phases, vault
    unlock/lock split, model checksum vs. load split, time to first token, document
    ingestion phases) to logs/perf.log. Off by default; no file is created without the
    variable. Records the cold-load figures docs/model-benchmarks.md §11.4 lists as
    still missing. See docs/benchmark.md "Perf marks".
  • Model checksum passes are visible now (issue #106). Verifying a model's
    multi-GB file — which can take minutes from a slow USB stick and used to run
    with no trace anywhere — now writes one diagnostics-log line per real
    verification (which model, how many bytes, how long), plus checksum_start /
    checksum_done marks in the opt-in perf log. Overlapping verifications of the
    same file (for example a start racing an AI Model screen visit) also no longer
    each read the whole file — they share one pass.
  • Honest progress while a model starts (issue #107). Once the app has
    measured your drive's real read speed, the Chat screen's "your model is
    starting" panel shows the model file size and an approximate percentage read
    instead of an indefinite spinner. Fresh installs (no measurement yet) keep the
    plain message rather than showing a made-up number.
  • Cold model starts are much faster on slow drives (issue #114). While a
    model loads, the app now also reads the model file front-to-back in the
    background, which lines the operating system's file cache up with what the
    load is about to need. Measured on a 16 GB laptop: a 6.65 GB model from a slow
    USB stick started in 5:46 instead of 11:20 (−49%), and from a portable SSD in
    10 s instead of 15 s (−36%). Starts that are already fast stay as they are —
    the helper is skipped when the file was just verified (its content is already
    in the cache) and stops the moment the load finishes or is cancelled.
  • One-click skill offers on document answers (issue #80). When a document
    answer cannot serve the shape you asked for — "categorize the transactions
    and sum per category" answered by an engine that can only list values — the
    answer now carries a one-click "Run "Bank Statement Analysis" for this
    question" action instead of only a prose pointer; clicking re-answers the
    same question with the skill (nothing ever runs without the click). On the
    residue the deterministic router provably cannot classify, a bounded,
    grammar-constrained on-device model pick (marked "Suggested by the local
    model") may fill the same offer; ordinary questions keep their zero-model-call
    routing, and every failure quietly falls back to today's answer.
  • Export your original files from the workspace (issue #90). Every imported
    document's ⋯ menu now offers Export original file, which saves the stored
    original — PDF, Word, recording, photo, any format — byte-for-byte to a
    location you choose, even when the file you imported it from is long gone.
    Imported documents previously had no export action at all (only generated
    documents did). The export shows the encryption-boundary warning first (the
    saved copy is not protected by your workspace password), writes atomically,
    and records only the document id in the activity log. One document per export
    for now; bulk export is a possible follow-up.

Changed

  • Qwen3.8 27B is the new recommended chat model for 24 GB and ≥32 GB machines
    (owner ratification 2026-08-16; docs/model-benchmarks.md §9.4). Three new
    catalog entries (Q4_K_M, Q5_K_M, and the selectable Q6_K quality ceiling for
    24 GB GPUs), all license-reviewed with verified upstream hashes. The Qwen3.6
    27B pair stays in the catalog, still ranked and selectable.
  • Every screen shares one content width (issues #166/#171). All screens —
    Home, Documents, Translate, Images, AI model, Skills, Settings — now use the
    same centred content width with symmetric side gutters, so pages no longer
    read as differently sized; only the chat keeps its full-bleed layout with the
    centred transcript. Explanatory texts inside are capped at a readable line
    length, so the wider pages don't make help texts harder to read. The centred
    content also no longer shifts sideways by the scrollbar's width when
    switching...
Read more

v0.1.56

Choose a tag to compare

@github-actions github-actions released this 10 Aug 00:22
e834234

Which file do I need?

The downloads are at the very bottom of this page, in the Assets list underneath all these notes. Scroll all the way down, and click Assets if the list is folded. You need exactly one file:

Your computer Download this ONE file
Windows 10/11 HilbertRaum-0.1.56-portable.exe
macOS (Apple Silicon) HilbertRaum-0.1.56-mac-arm64.app.zip
Linux HilbertRaum-0.1.56.AppImage

The other files are NOT the app: the llama-runtime-*.zip files carry the AI engine for offline/air-gapped installs only (normally the app offers that same download built in), SHA256SUMS.txt is for optional download verification, and the two Source code archives are the source for developers, not a runnable app.

This download is the app only — a working chat needs two more downloads, both offered on the AI Model screen inside the app: the AI engine (llama.cpp — the screen shows an install banner until it is present; without it, models run in demo mode with simulated answers) and an AI model of your choice. Every download asks first and is SHA-256-verified. Offline installs: the llama-runtime-*.zip assets below hold the same verified engine build — unzip into the runtime/llama.cpp/ folder on the drive. Repo users can provision a drive up front instead (README step 2). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.

Added

  • Local chat — a llama.cpp runtime running GGUF models entirely on-device
    (CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
    Granite) and an on-machine benchmark that recommends the best-fit model for
    available RAM. A built-in demo mode runs the whole UI with no model files
    and no network.
  • Document Q&A with citations — import PDF / Word / text, ask questions, and
    get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
    reranker, scoped by collections.
  • Image understanding — ask questions about a picture with a local vision
    model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
    and deletable — per-entry and total sizes are shown, a confirmed "Clear
    history" removes everything at once (issue #122), and the history's disk
    footprint appears in Settings → Diagnostics.
  • Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
    on-device OCR ("Make searchable") on scanned pages (bundled German + English
    language files; no cloud OCR).
  • Document tasks & skills — summarize, translate, and compare documents;
    install reusable skills for structured extraction (bank statements,
    invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe).
  • Evidence packs (review mode) — review a document-grounded answer block by
    block against its sources, record explicit decisions and notes, and export the
    review as a self-contained evidence pack in HTML or PDF — generated
    locally and offline, with honest coverage, freshness, and limitation notes. A
    pack supports human review; it is not a correctness certification.
  • Translate view + dedicated translation model (TranslateGemma) — a top-level
    Translate screen for live text translation and drag-and-drop document
    translation across 51 languages, source and target (the model's full
    production tier — from German and English to Arabic, Chinese, Swahili, and
    Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
    sidecar (downloaded on demand behind the license acknowledgement; not
    bundled) — never the chat model; document translations materialize as
    searchable, exportable local documents. GPU-accelerated when the machine
    allows it, with an automatic CPU fallback. Calibrated against the real model
    (per-language round-trip evidence + measured tokenizer weights) so a window
    can only over-chunk, never overflow.
  • Encrypted, portable workspace — an optional password-encrypted workspace
    (AES-256-GCM with Argon2id key derivation) covering the database, imported-document
    copies, and the diagnostics log; keep models plus the workspace on an external
    drive and move between laptops.
  • Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
    context-isolated renderer; a strict Content-Security-Policy; deny-by-default
    renderer permissions; an offline guard that trips on any non-loopback connection
    attempt; and a tamper-evident local audit log (ids/counts only, never content).
  • Cross-platform & distribution — Windows-first, with macOS and Linux supported
    in the architecture; portable / preconfigured-drive distribution via
    scripts/build-commercial-drive.*.
  • Standard project docs — this CHANGELOG.md and a
    CODE_OF_CONDUCT.md.
  • Windows engine build attached to releases (issue #102 item 1). Each release
    now carries llama-runtime-win-x64.zip beside the existing mac Metal zip: the
    same SHA-256-verified llama.cpp build the in-app installer fetches, packaged so
    it unzips straight into the drive's runtime/llama.cpp/ folder — an offline /
    air-gapped install path that needs no repo scripts.
  • Opt-in perf mark log for measurement runs: with HILBERTRAUM_PERF_LOG=1 set in
    the environment, the app appends timestamped timing marks (startup phases, vault
    unlock/lock split, model checksum vs. load split, time to first token, document
    ingestion phases) to logs/perf.log. Off by default; no file is created without the
    variable. Records the cold-load figures docs/model-benchmarks.md §11.4 lists as
    still missing. See docs/benchmark.md "Perf marks".
  • Model checksum passes are visible now (issue #106). Verifying a model's
    multi-GB file — which can take minutes from a slow USB stick and used to run
    with no trace anywhere — now writes one diagnostics-log line per real
    verification (which model, how many bytes, how long), plus checksum_start /
    checksum_done marks in the opt-in perf log. Overlapping verifications of the
    same file (for example a start racing an AI Model screen visit) also no longer
    each read the whole file — they share one pass.
  • Honest progress while a model starts (issue #107). Once the app has
    measured your drive's real read speed, the Chat screen's "your model is
    starting" panel shows the model file size and an approximate percentage read
    instead of an indefinite spinner. Fresh installs (no measurement yet) keep the
    plain message rather than showing a made-up number.
  • Cold model starts are much faster on slow drives (issue #114). While a
    model loads, the app now also reads the model file front-to-back in the
    background, which lines the operating system's file cache up with what the
    load is about to need. Measured on a 16 GB laptop: a 6.65 GB model from a slow
    USB stick started in 5:46 instead of 11:20 (−49%), and from a portable SSD in
    10 s instead of 15 s (−36%). Starts that are already fast stay as they are —
    the helper is skipped when the file was just verified (its content is already
    in the cache) and stops the moment the load finishes or is cancelled.
  • One-click skill offers on document answers (issue #80). When a document
    answer cannot serve the shape you asked for — "categorize the transactions
    and sum per category" answered by an engine that can only list values — the
    answer now carries a one-click "Run "Bank Statement Analysis" for this
    question" action instead of only a prose pointer; clicking re-answers the
    same question with the skill (nothing ever runs without the click). On the
    residue the deterministic router provably cannot classify, a bounded,
    grammar-constrained on-device model pick (marked "Suggested by the local
    model") may fill the same offer; ordinary questions keep their zero-model-call
    routing, and every failure quietly falls back to today's answer.
  • Export your original files from the workspace (issue #90). Every imported
    document's ⋯ menu now offers Export original file, which saves the stored
    original — PDF, Word, recording, photo, any format — byte-for-byte to a
    location you choose, even when the file you imported it from is long gone.
    Imported documents previously had no export action at all (only generated
    documents did). The export shows the encryption-boundary warning first (the
    saved copy is not protected by your workspace password), writes atomically,
    and records only the document id in the activity log. One document per export
    for now; bulk export is a possible follow-up.

Changed

  • Every screen shares one content width (issues #166/#171). All screens —
    Home, Documents, Translate, Images, AI model, Skills, Settings — now use the
    same centred content width with symmetric side gutters, so pages no longer
    read as differently sized; only the chat keeps its full-bleed layout with the
    centred transcript. Explanatory texts inside are capped at a readable line
    length, so the wider pages don't make help texts harder to read. The centred
    content also no longer shifts sideways by the scrollbar's width when
    switching between scrolling and non-scrolling pages.
  • Document translations pack fuller windows on compound-heavy languages
    (issue #165). The window planner now budgets and fills in the same unit
    (words), so German/Czech-style prose no longer splits into 1.5–2.5× more
    windows than needed — fewer model calls and materially less wall-clock on
    exactly the languages whose windows decode sl...
Read more

v0.1.55

Choose a tag to compare

@github-actions github-actions released this 23 Jul 15:51
bbf26ad

This download is the app only — the AI models are fetched separately (README step 2 / the in-app downloader). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.

Added

  • Local chat — a llama.cpp runtime running GGUF models entirely on-device
    (CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
    Granite) and an on-machine benchmark that recommends the best-fit model for
    available RAM. A built-in demo mode runs the whole UI with no model files
    and no network.
  • Document Q&A with citations — import PDF / Word / text, ask questions, and
    get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
    reranker, scoped by collections.
  • Image understanding — ask questions about a picture with a local vision
    model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
    and deletable.
  • Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
    on-device OCR ("Make searchable") on scanned pages (bundled German + English
    language files; no cloud OCR).
  • Document tasks & skills — summarize, translate, and compare documents;
    install reusable skills for structured extraction (bank statements,
    invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe).
  • Evidence packs (review mode) — review a document-grounded answer block by
    block against its sources, record explicit decisions and notes, and export the
    review as a self-contained evidence pack in HTML or PDF — generated
    locally and offline, with honest coverage, freshness, and limitation notes. A
    pack supports human review; it is not a correctness certification.
  • Translate view + dedicated translation model (TranslateGemma) — a top-level
    Translate screen for live text translation and drag-and-drop document
    translation across 51 languages, source and target (the model's full
    production tier — from German and English to Arabic, Chinese, Swahili, and
    Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
    sidecar (downloaded on demand behind the license acknowledgement; not
    bundled) — never the chat model; document translations materialize as
    searchable, exportable local documents. GPU-accelerated when the machine
    allows it, with an automatic CPU fallback. Calibrated against the real model
    (per-language round-trip evidence + measured tokenizer weights) so a window
    can only over-chunk, never overflow.
  • Encrypted, portable workspace — an optional password-encrypted workspace
    (AES-256-GCM with Argon2id key derivation) covering the database, imported-document
    copies, and the diagnostics log; keep models plus the workspace on an external
    drive and move between laptops.
  • Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
    context-isolated renderer; a strict Content-Security-Policy; deny-by-default
    renderer permissions; an offline guard that trips on any non-loopback connection
    attempt; and a tamper-evident local audit log (ids/counts only, never content).
  • Cross-platform & distribution — Windows-first, with macOS and Linux supported
    in the architecture; portable / preconfigured-drive distribution via
    scripts/build-commercial-drive.*.
  • Standard project docs — this CHANGELOG.md and a
    CODE_OF_CONDUCT.md.

Changed

  • Deep-index extraction is more reliable under reasoning-prone models — the
    "Build deep index" structured-extract pass now grammar-constrains the model's
    reply (the same JSON-schema mechanism the bank-statement categorizer uses), so
    sections can no longer come back unreadable because the model answered in
    prose or code fences; "unparsed" sections in listing answers should now be
    rare, and the existing retry/salvage safety net is unchanged (wave STR-1; see
    the docs/architecture.md "Skills & tools architecture review (2026-07-19) —
    design record").

Fixed

  • Scanned-PDF OCR is startable again from the Documents row (it had become
    unreachable after a row-actions refactor): "Make searchable (OCR)" is an inline
    button on the scan's row, already-recognized PDFs can be re-run via
    "Read again (OCR)", Translate now explains scanned PDFs (make searchable
    first) instead of calling them unsupported, and progress is honest through the
    final "Finishing" step. Packaged builds no longer carry the dev-only localhost
    CSP relaxation in their HTML meta tags (wave OCR-R, PR #75; see the
    docs/architecture.md "OCR audit (2026-07-18) — remediation ledger").
  • npm run dev no longer 500s on the first page load — a false "no CSP meta
    tag" throw during dev serve (the guard mis-read a deliberate, byte-identical
    no-op rewrite as a missing tag) is fixed; packaged builds were never affected
    (wave DEP-1, PR #77).

Security

  • Post-DEP-1 advisory batch cleared (wave DEP-2, 2026-07-23) — five transitive
    dev/build-tooling packages patched lockfile-only, semver-compatible, no manifest
    changes: node-tar 7.5.21 (decompression/parse DoS CRITICAL + infinite-loop,
    PAX-crash and NUL-byte advisories), js-yaml 4.3.0 (merge-key quadratic CPU;
    electron-builder's copy — the app's own manifest parser is the separate yaml
    package and was never affected), fast-uri 3.1.4 (host confusion ×2),
    brace-expansion 1.1.16 / 2.1.2 / 5.0.8 (exponential-time {} expansion DoS,
    CVE-2026-13149 — Dependabot alert #56), and DOMPurify 3.4.12 (the one
    production-scope member: streamdown → mermaid ships it in the renderer;
    CUSTOM_ELEMENT_HANDLING sanitizer bypass, low severity, config not used by
    mermaid). npm audit: 0 vulnerabilities again.
  • All critical- and high-severity Dependabot alerts cleared (wave DEP-1, PR #77)
    — Vitest 3.2.6 (CVE-2026-47429, UI-server arbitrary file read/execute), Electron
    39.8.10 (command-line switch injection, four use-after-free classes,
    permission-origin confusion, header injection), Vite 6.4.3 + electron-vite 3.1.0
    (server.fs.deny bypass on Windows, path traversal), form-data 4.0.6 (CRLF
    injection, CVE-2026-12143), undici 7.28.0 + 6.27.0 (TLS-bypass and cross-origin
    routing via a SOCKS5 proxy, header injection), and esbuild 0.25.12 (dev-server
    CORS). npm audit now reports 0 vulnerabilities. Packaged-build security was
    re-verified on the new Electron 39 runtime: the strict Content-Security-Policy
    response header still attaches and enforces on file:// in both windows (see
    docs/architecture.md "Dependency remediation — design record (wave DEP-1,
    PR #77)").

v0.1.54

Choose a tag to compare

@github-actions github-actions released this 19 Jul 10:26
e6af01b

This download is the app only — the AI models are fetched separately (README step 2 / the in-app downloader). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.

Added

  • Local chat — a llama.cpp runtime running GGUF models entirely on-device
    (CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
    Granite) and an on-machine benchmark that recommends the best-fit model for
    available RAM. A built-in demo mode runs the whole UI with no model files
    and no network.
  • Document Q&A with citations — import PDF / Word / text, ask questions, and
    get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
    reranker, scoped by collections.
  • Image understanding — ask questions about a picture with a local vision
    model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
    and deletable.
  • Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
    on-device OCR ("Make searchable") on scanned pages (bundled German + English
    language files; no cloud OCR).
  • Document tasks & skills — summarize, translate, and compare documents;
    install reusable skills for structured extraction (bank statements,
    invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe).
  • Evidence packs (review mode) — review a document-grounded answer block by
    block against its sources, record explicit decisions and notes, and export the
    review as a self-contained evidence pack in HTML or PDF — generated
    locally and offline, with honest coverage, freshness, and limitation notes. A
    pack supports human review; it is not a correctness certification.
  • Translate view + dedicated translation model (TranslateGemma) — a top-level
    Translate screen for live text translation and drag-and-drop document
    translation across 51 languages, source and target (the model's full
    production tier — from German and English to Arabic, Chinese, Swahili, and
    Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
    sidecar (downloaded on demand behind the license acknowledgement; not
    bundled) — never the chat model; document translations materialize as
    searchable, exportable local documents. GPU-accelerated when the machine
    allows it, with an automatic CPU fallback. Calibrated against the real model
    (per-language round-trip evidence + measured tokenizer weights) so a window
    can only over-chunk, never overflow.
  • Encrypted, portable workspace — an optional password-encrypted workspace
    (AES-256-GCM with Argon2id key derivation) covering the database, imported-document
    copies, and the diagnostics log; keep models plus the workspace on an external
    drive and move between laptops.
  • Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
    context-isolated renderer; a strict Content-Security-Policy; deny-by-default
    renderer permissions; an offline guard that trips on any non-loopback connection
    attempt; and a tamper-evident local audit log (ids/counts only, never content).
  • Cross-platform & distribution — Windows-first, with macOS and Linux supported
    in the architecture; portable / preconfigured-drive distribution via
    scripts/build-commercial-drive.*.
  • Standard project docs — this CHANGELOG.md and a
    CODE_OF_CONDUCT.md.

Fixed

  • Scanned-PDF OCR is startable again from the Documents row (it had become
    unreachable after a row-actions refactor): "Make searchable (OCR)" is an inline
    button on the scan's row, already-recognized PDFs can be re-run via
    "Read again (OCR)", Translate now explains scanned PDFs (make searchable
    first) instead of calling them unsupported, and progress is honest through the
    final "Finishing" step. Packaged builds no longer carry the dev-only localhost
    CSP relaxation in their HTML meta tags (wave OCR-R, PR #75; see the
    docs/architecture.md "OCR audit (2026-07-18) — remediation ledger").
  • npm run dev no longer 500s on the first page load — a false "no CSP meta
    tag" throw during dev serve (the guard mis-read a deliberate, byte-identical
    no-op rewrite as a missing tag) is fixed; packaged builds were never affected
    (wave DEP-1, PR #77).

Security

  • All critical- and high-severity Dependabot alerts cleared (wave DEP-1, PR #77)
    — Vitest 3.2.6 (CVE-2026-47429, UI-server arbitrary file read/execute), Electron
    39.8.10 (command-line switch injection, four use-after-free classes,
    permission-origin confusion, header injection), Vite 6.4.3 + electron-vite 3.1.0
    (server.fs.deny bypass on Windows, path traversal), form-data 4.0.6 (CRLF
    injection, CVE-2026-12143), undici 7.28.0 + 6.27.0 (TLS-bypass and cross-origin
    routing via a SOCKS5 proxy, header injection), and esbuild 0.25.12 (dev-server
    CORS). npm audit now reports 0 vulnerabilities. Packaged-build security was
    re-verified on the new Electron 39 runtime: the strict Content-Security-Policy
    response header still attaches and enforces on file:// in both windows (see
    docs/architecture.md "Dependency remediation — design record (wave DEP-1,
    PR #77)").

v0.1.53

Choose a tag to compare

@github-actions github-actions released this 18 Jul 15:23
e29b510

This download is the app only — the AI models are fetched separately (README step 2 / the in-app downloader). Unsigned build: see the SmartScreen/Gatekeeper notes in docs/troubleshooting.md.

The accumulated, feature-complete MVP — plus the GPU-acceleration,
retrieval-quality, UI-polish, and office-functionality waves — to be cut as the
first public release. Consciously-accepted gaps are tracked in
docs/known-limitations.md.

Added

  • Local chat — a llama.cpp runtime running GGUF models entirely on-device
    (CPU or GPU), with a curated open-weight catalog (Qwen3, Ministral, Gemma,
    Granite) and an on-machine benchmark that recommends the best-fit model for
    available RAM. A built-in demo mode runs the whole UI with no model files
    and no network.
  • Document Q&A with citations — import PDF / Word / text, ask questions, and
    get answers grounded in your files. Hybrid (vector + keyword) retrieval with a
    reranker, scoped by collections.
  • Image understanding — ask questions about a picture with a local vision
    model (Qwen2.5-VL); the analysis history is stored locally, encrypted at rest,
    and deletable.
  • Audio & voice — transcribe audio files (Whisper), dictate prompts, and run
    on-device OCR ("Make searchable") on scanned pages (bundled German + English
    language files; no cloud OCR).
  • Document tasks & skills — summarize, translate, and compare documents;
    install reusable skills for structured extraction (bank statements,
    invoices, meeting minutes, contract briefs, deadlines, redaction / share-safe).
  • Evidence packs (review mode) — review a document-grounded answer block by
    block against its sources, record explicit decisions and notes, and export the
    review as a self-contained evidence pack in HTML or PDF — generated
    locally and offline, with honest coverage, freshness, and limitation notes. A
    pack supports human review; it is not a correctness certification.
  • Translate view + dedicated translation model (TranslateGemma) — a top-level
    Translate screen for live text translation and drag-and-drop document
    translation across 51 languages, source and target (the model's full
    production tier — from German and English to Arabic, Chinese, Swahili, and
    Vietnamese). Translation runs on a dedicated on-device TranslateGemma 12B
    sidecar (downloaded on demand behind the license acknowledgement; not
    bundled) — never the chat model; document translations materialize as
    searchable, exportable local documents. GPU-accelerated when the machine
    allows it, with an automatic CPU fallback. Calibrated against the real model
    (per-language round-trip evidence + measured tokenizer weights) so a window
    can only over-chunk, never overflow.
  • Encrypted, portable workspace — an optional password-encrypted workspace
    (AES-256-GCM with Argon2id key derivation) covering the database, imported-document
    copies, and the diagnostics log; keep models plus the workspace on an external
    drive and move between laptops.
  • Privacy & security posture — no cloud, telemetry, or analytics; a sandboxed,
    context-isolated renderer; a strict Content-Security-Policy; deny-by-default
    renderer permissions; an offline guard that trips on any non-loopback connection
    attempt; and a tamper-evident local audit log (ids/counts only, never content).
  • Cross-platform & distribution — Windows-first, with macOS and Linux supported
    in the architecture; portable / preconfigured-drive distribution via
    scripts/build-commercial-drive.*.
  • Standard project docs — this CHANGELOG.md and a
    CODE_OF_CONDUCT.md.