Repository navigation
Releases: Igna-Mendez/Open-Local-Audio-Scribe-OLAS
Release list
Release 1.1
What's new in 1.1
- A model choice, and two modes. First run asks whether you want
Normal (more accurate) or Potato (lighter on your CPU — good for a busy
or older machine). Change it any time from Options. - Steadier under load. Transcripts no longer fall behind on long stretches
of talking, and the app no longer hogs the machine while it works. - Starts up faster, and downloads what it needs automatically.
- Update check works now — it tells you when a newer version is out.
Still two languages side by side, still entirely offline.
OLAS — Open Local Audio Scribe (Windows) — version 1.1
Real-time local speech-to-text for Windows. Two language panes side by side,
one model each, running entirely on your machine. No cloud, no telemetry, no
account, no API keys, no network at runtime.
Built for meetings, interviews, lectures and interpretation — the situations
where a transcript needs to keep up with people talking, and where the audio
should never leave the machine. Captures what the machine is playing via
WASAPI loopback, so it works on calls, videos and streams without a virtual
cable.
Accuracy is the design goal. Where a trade-off exists between staying
current and keeping every word, this program keeps the word.
Note on how this was built. OLAS was heavily vibe-coded: roughly 95% of
the code was written by different agentic AI models, with a human directing
the design, testing on real hardware and deciding what shipped. The
measurements in PATCHNOTES.md exist because the AI-written
parts got things confidently wrong more than once, and only measurement
caught it. Treat the code accordingly: it works, but it has not had a
conventional human review.
Download and run
- Open the
Releases page. - Download the latest
OLAS-win64-1.1.x.zip. - Extract it anywhere.
- Double-click
OLAS.bat.
The first launch asks which English model to load (see Modes); after
that it starts straight up. The zip contains everything — the executable, the
runtime DLL, and all three models. Nothing else to install.
Windows will warn you the first time. The executable is not code-signed,
so SmartScreen shows "Windows protected your PC". Click More info →
Run anyway. This is normal for an unsigned open-source build and not a sign
of anything wrong with the file.
Requirements
- Windows 10 21H2 or newer (Windows 11 recommended)
- x64 CPU with AVX2 — Intel Haswell (2013) or AMD Excavator (2015) and newer
- 4+ logical CPU threads recommended
- ~800 MB free disk
No GPU is used or required.
What it does
- Two languages at once — English and Spanish, each in its own pane with
independent Start/Stop, collapse and decouple-to-window controls - Two modes — Normal (more accurate) and Potato (ultralight), chosen on
first run and changeable from Options - Focus mode — hides the toolbar and pane headers so the transcript fills
the window - Live transcript file — written line by line to
olas-moonshine-notes.txt
beside the executable - Options popup — capture device, English model, zoom, timestamps,
light/dark theme, auto-scroll - Update check — asks GitHub whether a newer release exists
- Diagnostics —
-vwrites per-line latency toolas-debug.log,--stats
prints inference diagnostics on exit
Modes
Which mode runs is decided by the English model. Spanish is always Small
Streaming: no Medium Spanish model exists.
| mode | English | Spanish | cores | character |
|---|---|---|---|---|
| Normal (default) | Medium Streaming | Small Streaming | 3 + 1 | more accurate |
| Potato | Small Streaming | Small Streaming | 1 + 1 | ultralight, lower CPU, slightly less accurate |
The English model is chosen on first launch and remembered in olas-model.txt
beside the executable. To change it later: Options → English model, then
restart — the new model is loaded on the next launch.
How it works
WASAPI loopback (miniaudio)
| 16 kHz mono s16, 50 ms chunks
v
capture.c ring buffer
| capture_read_chunk()
v
capture thread
| Engine::feed() -> AudioQueue per language (lossless, non-blocking)
v
worker thread per language
| Moonshine Transcriber (streaming)
v
listener -> Win32 RichEdit, and to olas-moonshine-notes.txt
Two models run in parallel, each processing the same audio stream. Both
transcribe everything; you read the pane for the language being spoken. This is
simpler and more robust than language detection, at the cost of the non-target
pane producing nonsense — a known limitation, discussed in
PATCHNOTES.md.
Each language runs on its own worker thread with its own Transcriber, on its
own CPU core set. The core budget and per-model split are derived from the
detected hardware at startup, so a 2-core laptop and a 16-core desktop each get
a sensible share without configuration.
Source layout
src/
main_win.cpp entry point, CLI, capture thread
moonshine_engine.* Engine, worker slots, Streaming/NonStreaming workers
resource_plan.* CPU detection, per-model core split, affinity
model_choice.* persisted English model selection
win32_ui.* panes, toolbar overlay, buttons, popup menu
capture.c / .h WASAPI loopback + ring buffer
miniaudio.h vendored miniaudio 0.11
update_check.* GitHub release check
transcript_sink.h engine <-> UI abstraction
tools/
fetch-streaming-models.ps1
cmake/
copy-models.cmake bundles models into the release zip
Configuration
olas-win.conf holds the transcription parameters:
[general]
vad_threshold = 0.5 # speech/silence threshold
vad_max_segment_duration = 12 # longest single line, seconds
transcription_interval = 1.0 # how often the decoder re-runsDelete it and built-in defaults apply. Bad values and unknown keys warn on
stderr; they never stop the program starting.
Threading, affinity and the model choice are not configured here — they come
from the hardware probe and from olas-model.txt.
Command line
| Flag | Description | Default |
|---|---|---|
-l, --language CODE[,CODE] |
Language codes | en,es |
-m, --model PATH[,PATH] |
Model directory per language | resolved from the arch |
-a, --arch N[,N] |
0=Tiny, 1=Base, 2=TinyStreaming, 4=SmallStreaming, 5=MediumStreaming | from olas-model.txt |
-q, --chunk-ms MS |
Capture chunk (20..1000) | 50 |
-v, --verbose |
Write olas-debug.log |
off |
-c, --config PATH |
Transcription parameters file | olas-win.conf |
--no-update-check |
Skip the release check | off |
--stats |
Print inference diagnostics on exit | off |
-h, --help |
Show help |
3=BaseStreaming is rejected — declared in Moonshine's C API for
forward-compatibility but not supported.
Built with
- Moonshine Voice — the
speech-to-text models and C++ runtime. Small and Medium Streaming
architectures, MIT licensed. - Win32 / RichEdit — the user interface, with a custom button class for
theme-aware controls. - miniaudio — WASAPI loopback
capture, vendored single-header. - ONNX Runtime — inference, shipped inside Moonshine's SDK.
Building from source
Prerequisites: Visual Studio 2022 Build Tools with "Desktop development with
C++", CMake ≥ 3.20, Git, PowerShell 5.1+, and optionally Ninja.
git clone https://github.com/Igna-Mendez/Open-Local-Audio-Scribe-OLAS.git
cd Open-Local-Audio-Scribe-OLAS
powershell -ExecutionPolicy Bypass -File setup.ps1
mkdir build
cd build
cmake -G Ninja -DCMAKE_BUILD_TYPE=Release ..
cmake --build . --config Release -jPowerShell blocks scripts by default, hence -ExecutionPolicy Bypass -File.
setup.ps1 checks the toolchain, fetches all three models into models\, and
prints the exact configure command.
The first configure downloads the Moonshine SDK — a prebuilt release
archive (~26 MB) containing the headers, the static libraries and
onnxruntime.dll. Nothing is compiled from source, so configure takes a minute
or two rather than the 5–15 minutes a source build of ONNX Runtime would need.
It follows Moonshine's latest release and stays current: the configure
records which release it fetched, checks for a newer one each time, and
re-downloads when there is one.
| option | effect |
|---|---|
-DMOONSHINE_VERSION=v0.1.5 |
pin a specific release instead of latest |
-DMOONSHINE_REFRESH=1 |
force a re-download now |
-DMOONSHINE_SDK_DIR=<dir> |
use an SDK already on disk |
Deleting build\moonshine-sdk\ also forces a fresh download. Offline, or if
the release check fails, the cached SDK is kept rather than failing the
configure.
Debug builds are rejected at configure time — they are roughly 5× slower and
never useful for judging transcription quality.
Making a release zip
From build\:
cmake --build . --config Release --target zipCreates build\dist\OLAS-win64-<version>.zip, containing the executable, the
runtime DLL, all three models, the launcher and the docs — everything an end
user needs. Run setup.ps1 first so models\ is populated; the target warns
rather than failing if it is empty, but the zip will not run without them.
Cross-compiling from Linux with mingw-w64 is also supported via
mingw-toolchain.cmake.
Known limitations
- The non-target pane produces nonsense. Both models h...
Olas v0.6 Windows Release
OLAS 0.6.0 — Small Streaming release
First public Windows release of OLAS. Runs Moonshine C++ locally with WASAPI loopback capture. No cloud, no telemetry, no external dependencies beyond the bundled ONNX Runtime.
Highlights
- Small Streaming models for English and Spanish — MIT-licensed, high accuracy, low latency
- Two languages side by side with independent Start/Stop and Detach-to-window controls
- Focus mode — collapse the toolbar and pane headers so the transcript fills the window
- Auto-scroll with proper user-scroll detection
- Options popup with device picker, zoom presets, timestamp toggle, dark/light theme, auto-scroll reset
- Custom rounded buttons that respect the dark theme
- Bounded resources — 6 s audio queue, 4000-line transcript cap, one core per language
Fixes since the pre-release build
- Restored normal thread priority for the inference worker
- Restored transcriber error reporting and audio queue overflow warnings
- Fixed the
default_model_dirpath separator on Windows - Fixed pane layout on detach/reattach
- Fixed toolbar overlay z-order after pane reparenting
- Custom buttons repaint correctly on theme switch
- Detached panes can now be minimized and restored from the taskbar
- Main window has no artificial minimum size
- Restore button handles destroyed/orphaned child windows
Requirements
- Windows 10 21H2 or newer (Windows 11 recommended)
- x64 CPU with AVX2 (Intel Haswell / AMD Excavator or newer)
- ~500 MB free disk for the models
- 4+ logical CPU threads recommended
Installation
- Download
OLAS-win64-0.6.0.zipbelow. - Extract anywhere.
- Double-click
OLAS.bat.
First launch takes a few seconds while the models load.
Building from source
See README.md. Requires VS 2022 x64 toolchain, CMake ≥ 3.20, and PowerShell 5.1+.
V0.5 Open Local Audio Scribe - Olas Windows x64
Major performance improvements in this version reworked constricting params to lead that lead to a choke point in the transcription in the past build infering whether to rework whole recapture section for better performance in future updates
Requirements
Windows 10 1809 or newer (Windows 11 recommended)
A working audio output device (WASAPI)
Quick start
Double-click OLAS.bat
Play audio through your default output device
The window will transcribe what you hear
What it does
OLAS listens to whatever your system plays (WASAPI loopback) and
transcribes it locally using Moonshine. No internet needed after
first launch. Supports English and Spanish by default.
Troubleshooting
"Could not initialize audio"
Your system has no WASAPI output device. Check Windows
Settings -> System -> Sound.
"Failed to initialise Moonshine"
Model files are missing. Ensure the models\ folder sits next
to olas_win.exe with base-en\ and base-es\ subfolders.
Nothing appears when audio is playing
Make sure audio is playing through the DEFAULT output device.
Loopback captures that device, not the microphone.
Project page