Skip to content

Releases: Igna-Mendez/Open-Local-Audio-Scribe-OLAS

Release 1.1

Choose a tag to compare

@Igna-Mendez Igna-Mendez released this 08 Oct 19:35

What's new in 1.1

  • A model choice, and two modes. First run asks whether you want
    Normal (more accurate) or Potato (lighter on your CPU — good for a busy
    or older machine). Change it any time from Options.
  • Steadier under load. Transcripts no longer fall behind on long stretches
    of talking, and the app no longer hogs the machine while it works.
  • Starts up faster, and downloads what it needs automatically.
  • Update check works now — it tells you when a newer version is out.

Still two languages side by side, still entirely offline.

OLAS — Open Local Audio Scribe (Windows) — version 1.1

Real-time local speech-to-text for Windows. Two language panes side by side,
one model each, running entirely on your machine. No cloud, no telemetry, no
account, no API keys, no network at runtime.

Built for meetings, interviews, lectures and interpretation — the situations
where a transcript needs to keep up with people talking, and where the audio
should never leave the machine. Captures what the machine is playing via
WASAPI loopback, so it works on calls, videos and streams without a virtual
cable.

Accuracy is the design goal. Where a trade-off exists between staying
current and keeping every word, this program keeps the word.

Note on how this was built. OLAS was heavily vibe-coded: roughly 95% of
the code was written by different agentic AI models, with a human directing
the design, testing on real hardware and deciding what shipped. The
measurements in PATCHNOTES.md exist because the AI-written
parts got things confidently wrong more than once, and only measurement
caught it. Treat the code accordingly: it works, but it has not had a
conventional human review.


Download and run

  1. Open the
    Releases page.
  2. Download the latest OLAS-win64-1.1.x.zip.
  3. Extract it anywhere.
  4. Double-click OLAS.bat.

The first launch asks which English model to load (see Modes); after
that it starts straight up. The zip contains everything — the executable, the
runtime DLL, and all three models. Nothing else to install.

Windows will warn you the first time. The executable is not code-signed,
so SmartScreen shows "Windows protected your PC". Click More info →
Run anyway
. This is normal for an unsigned open-source build and not a sign
of anything wrong with the file.

Requirements

  • Windows 10 21H2 or newer (Windows 11 recommended)
  • x64 CPU with AVX2 — Intel Haswell (2013) or AMD Excavator (2015) and newer
  • 4+ logical CPU threads recommended
  • ~800 MB free disk

No GPU is used or required.

What it does

  • Two languages at once — English and Spanish, each in its own pane with
    independent Start/Stop, collapse and decouple-to-window controls
  • Two modes — Normal (more accurate) and Potato (ultralight), chosen on
    first run and changeable from Options
  • Focus mode — hides the toolbar and pane headers so the transcript fills
    the window
  • Live transcript file — written line by line to olas-moonshine-notes.txt
    beside the executable
  • Options popup — capture device, English model, zoom, timestamps,
    light/dark theme, auto-scroll
  • Update check — asks GitHub whether a newer release exists
  • Diagnostics — -v writes per-line latency to olas-debug.log, --stats
    prints inference diagnostics on exit

Modes

Which mode runs is decided by the English model. Spanish is always Small
Streaming: no Medium Spanish model exists.

mode English Spanish cores character
Normal (default) Medium Streaming Small Streaming 3 + 1 more accurate
Potato Small Streaming Small Streaming 1 + 1 ultralight, lower CPU, slightly less accurate

The English model is chosen on first launch and remembered in olas-model.txt
beside the executable. To change it later: Options → English model, then
restart — the new model is loaded on the next launch.

How it works

WASAPI loopback (miniaudio)
        |  16 kHz mono s16, 50 ms chunks
        v
capture.c ring buffer
        |  capture_read_chunk()
        v
capture thread
        |  Engine::feed() -> AudioQueue per language (lossless, non-blocking)
        v
worker thread per language
        |  Moonshine Transcriber (streaming)
        v
listener -> Win32 RichEdit, and to olas-moonshine-notes.txt

Two models run in parallel, each processing the same audio stream. Both
transcribe everything; you read the pane for the language being spoken. This is
simpler and more robust than language detection, at the cost of the non-target
pane producing nonsense — a known limitation, discussed in
PATCHNOTES.md.

Each language runs on its own worker thread with its own Transcriber, on its
own CPU core set. The core budget and per-model split are derived from the
detected hardware at startup, so a 2-core laptop and a 16-core desktop each get
a sensible share without configuration.

Source layout

src/
    main_win.cpp         entry point, CLI, capture thread
    moonshine_engine.*   Engine, worker slots, Streaming/NonStreaming workers
    resource_plan.*      CPU detection, per-model core split, affinity
    model_choice.*       persisted English model selection
    win32_ui.*           panes, toolbar overlay, buttons, popup menu
    capture.c / .h       WASAPI loopback + ring buffer
    miniaudio.h          vendored miniaudio 0.11
    update_check.*       GitHub release check
    transcript_sink.h    engine <-> UI abstraction
tools/
    fetch-streaming-models.ps1
cmake/
    copy-models.cmake    bundles models into the release zip

Configuration

olas-win.conf holds the transcription parameters:

[general]
vad_threshold = 0.5             # speech/silence threshold
vad_max_segment_duration = 12   # longest single line, seconds
transcription_interval = 1.0    # how often the decoder re-runs

Delete it and built-in defaults apply. Bad values and unknown keys warn on
stderr; they never stop the program starting.

Threading, affinity and the model choice are not configured here — they come
from the hardware probe and from olas-model.txt.

Command line

Flag Description Default
-l, --language CODE[,CODE] Language codes en,es
-m, --model PATH[,PATH] Model directory per language resolved from the arch
-a, --arch N[,N] 0=Tiny, 1=Base, 2=TinyStreaming, 4=SmallStreaming, 5=MediumStreaming from olas-model.txt
-q, --chunk-ms MS Capture chunk (20..1000) 50
-v, --verbose Write olas-debug.log off
-c, --config PATH Transcription parameters file olas-win.conf
--no-update-check Skip the release check off
--stats Print inference diagnostics on exit off
-h, --help Show help

3=BaseStreaming is rejected — declared in Moonshine's C API for
forward-compatibility but not supported.

Built with

  • Moonshine Voice — the
    speech-to-text models and C++ runtime. Small and Medium Streaming
    architectures, MIT licensed.
  • Win32 / RichEdit — the user interface, with a custom button class for
    theme-aware controls.
  • miniaudio — WASAPI loopback
    capture, vendored single-header.
  • ONNX Runtime — inference, shipped inside Moonshine's SDK.

Building from source

Prerequisites: Visual Studio 2022 Build Tools with "Desktop development with
C++", CMake ≥ 3.20, Git, PowerShell 5.1+, and optionally Ninja.

git clone https://github.com/Igna-Mendez/Open-Local-Audio-Scribe-OLAS.git
cd Open-Local-Audio-Scribe-OLAS
powershell -ExecutionPolicy Bypass -File setup.ps1

mkdir build
cd build
cmake -G Ninja -DCMAKE_BUILD_TYPE=Release ..
cmake --build . --config Release -j

PowerShell blocks scripts by default, hence -ExecutionPolicy Bypass -File.
setup.ps1 checks the toolchain, fetches all three models into models\, and
prints the exact configure command.

The first configure downloads the Moonshine SDK — a prebuilt release
archive (~26 MB) containing the headers, the static libraries and
onnxruntime.dll. Nothing is compiled from source, so configure takes a minute
or two rather than the 5–15 minutes a source build of ONNX Runtime would need.

It follows Moonshine's latest release and stays current: the configure
records which release it fetched, checks for a newer one each time, and
re-downloads when there is one.

option effect
-DMOONSHINE_VERSION=v0.1.5 pin a specific release instead of latest
-DMOONSHINE_REFRESH=1 force a re-download now
-DMOONSHINE_SDK_DIR=<dir> use an SDK already on disk

Deleting build\moonshine-sdk\ also forces a fresh download. Offline, or if
the release check fails, the cached SDK is kept rather than failing the
configure.

Debug builds are rejected at configure time — they are roughly 5× slower and
never useful for judging transcription quality.

Making a release zip

From build\:

cmake --build . --config Release --target zip

Creates build\dist\OLAS-win64-<version>.zip, containing the executable, the
runtime DLL, all three models, the launcher and the docs — everything an end
user needs. Run setup.ps1 first so models\ is populated; the target warns
rather than failing if it is empty, but the zip will not run without them.

Cross-compiling from Linux with mingw-w64 is also supported via
mingw-toolchain.cmake.

Known limitations

  • The non-target pane produces nonsense. Both models h...
Read more

Olas v0.6 Windows Release

Choose a tag to compare

@Igna-Mendez Igna-Mendez released this 28 Sep 18:22

OLAS 0.6.0 — Small Streaming release

First public Windows release of OLAS. Runs Moonshine C++ locally with WASAPI loopback capture. No cloud, no telemetry, no external dependencies beyond the bundled ONNX Runtime.

Highlights

  • Small Streaming models for English and Spanish — MIT-licensed, high accuracy, low latency
  • Two languages side by side with independent Start/Stop and Detach-to-window controls
  • Focus mode — collapse the toolbar and pane headers so the transcript fills the window
  • Auto-scroll with proper user-scroll detection
  • Options popup with device picker, zoom presets, timestamp toggle, dark/light theme, auto-scroll reset
  • Custom rounded buttons that respect the dark theme
  • Bounded resources — 6 s audio queue, 4000-line transcript cap, one core per language

Fixes since the pre-release build

  • Restored normal thread priority for the inference worker
  • Restored transcriber error reporting and audio queue overflow warnings
  • Fixed the default_model_dir path separator on Windows
  • Fixed pane layout on detach/reattach
  • Fixed toolbar overlay z-order after pane reparenting
  • Custom buttons repaint correctly on theme switch
  • Detached panes can now be minimized and restored from the taskbar
  • Main window has no artificial minimum size
  • Restore button handles destroyed/orphaned child windows

Requirements

  • Windows 10 21H2 or newer (Windows 11 recommended)
  • x64 CPU with AVX2 (Intel Haswell / AMD Excavator or newer)
  • ~500 MB free disk for the models
  • 4+ logical CPU threads recommended

Installation

  1. Download OLAS-win64-0.6.0.zip below.
  2. Extract anywhere.
  3. Double-click OLAS.bat.

First launch takes a few seconds while the models load.

Building from source

See README.md. Requires VS 2022 x64 toolchain, CMake ≥ 3.20, and PowerShell 5.1+.

V0.5 Open Local Audio Scribe - Olas Windows x64

Choose a tag to compare

@Igna-Mendez Igna-Mendez released this 28 Sep 03:33

Major performance improvements in this version reworked constricting params to lead that lead to a choke point in the transcription in the past build infering whether to rework whole recapture section for better performance in future updates

Requirements

Windows 10 1809 or newer (Windows 11 recommended)
A working audio output device (WASAPI)
Quick start

Double-click OLAS.bat
Play audio through your default output device
The window will transcribe what you hear

What it does

OLAS listens to whatever your system plays (WASAPI loopback) and
transcribes it locally using Moonshine. No internet needed after
first launch. Supports English and Spanish by default.
Troubleshooting

"Could not initialize audio"
Your system has no WASAPI output device. Check Windows
Settings -> System -> Sound.

"Failed to initialise Moonshine"
Model files are missing. Ensure the models\ folder sits next
to olas_win.exe with base-en\ and base-es\ subfolders.

Nothing appears when audio is playing
Make sure audio is playing through the DEFAULT output device.
Loopback captures that device, not the microphone.
Project page

https://github.com/Igna-Mendez/Olas