v1.0.1
EchoNotes v1.0.1
This release packages the work from the last development stretch into the first stable Dockerized worker release of EchoNotes.
Highlights
- Switched audio transcription from Whisper to WhisperX.
- Added persistent model loading so ASR models are loaded once and reused.
- Refactored the app into a queue-backed worker pool instead of processing inline in the watcher.
- Added transcript formatting, chunked summarization, and provider-based LLM support for Open WebUI, Ollama, OpenAI, Anthropic/Claude, and OpenRouter.
- Added Obsidian vault export with linked MP3 timestamp references and speaker-aware transcript output.
- Added WhisperX diarization support for
Speaker 1,Speaker 2, and so on. - Expanded audio ingest to common FFmpeg-readable formats and normalized them to MP3 before transcription.
- Hardened file ingestion against partial uploads, temporary files, encrypted PDFs, and OCR dependency failures.
- Added CPU and CUDA Docker image variants with mounted config, incoming, vault, and model-cache directories.
- Added warm-cache tooling for WhisperX models and documented the Docker build/runtime flow.
Docker Images
This release publishes:
robchartier/echonotes:1.0.1robchartier/echonotes:1.0.1-cuda12.8robchartier/echonotes:latestrobchartier/echonotes:latest-cuda12.8robchartier/echonotes:gpu
Operational Notes
- GPU deployments should generally use
worker_count: 1unless you have verified VRAM headroom for multiple concurrent WhisperX workers. - WhisperX diarization depends on a Hugging Face token and will fall back cleanly if diarization is not configured.
- Files copied into Windows-backed bind mounts may not always emit reliable filesystem events into Docker; EchoNotes now queues files already present at startup, but Linux-side writes remain the most reliable path.
QA Notes
Validated against the current Testing/ corpus on the CUDA image:
- Standard PDF, DOCX, TXT, FLAC, and WAV samples completed successfully.
- Encrypted PDFs were skipped gracefully without crashing the worker.
- Short audio files produced transcript, summary, Obsidian note, and vault copies correctly.
Known issues from QA:
- Very large audio can still hit GPU memory limits depending on model choice and VRAM.
- Very large text inputs can take long enough to block a single-worker queue under slow local LLMs.