Skip to content

v1.0.1

Choose a tag to compare

@github-actions github-actions released this 06 Mar 23:28
· 6 commits to main since this release

EchoNotes v1.0.1

This release packages the work from the last development stretch into the first stable Dockerized worker release of EchoNotes.

Highlights

  • Switched audio transcription from Whisper to WhisperX.
  • Added persistent model loading so ASR models are loaded once and reused.
  • Refactored the app into a queue-backed worker pool instead of processing inline in the watcher.
  • Added transcript formatting, chunked summarization, and provider-based LLM support for Open WebUI, Ollama, OpenAI, Anthropic/Claude, and OpenRouter.
  • Added Obsidian vault export with linked MP3 timestamp references and speaker-aware transcript output.
  • Added WhisperX diarization support for Speaker 1, Speaker 2, and so on.
  • Expanded audio ingest to common FFmpeg-readable formats and normalized them to MP3 before transcription.
  • Hardened file ingestion against partial uploads, temporary files, encrypted PDFs, and OCR dependency failures.
  • Added CPU and CUDA Docker image variants with mounted config, incoming, vault, and model-cache directories.
  • Added warm-cache tooling for WhisperX models and documented the Docker build/runtime flow.

Docker Images

This release publishes:

  • robchartier/echonotes:1.0.1
  • robchartier/echonotes:1.0.1-cuda12.8
  • robchartier/echonotes:latest
  • robchartier/echonotes:latest-cuda12.8
  • robchartier/echonotes:gpu

Operational Notes

  • GPU deployments should generally use worker_count: 1 unless you have verified VRAM headroom for multiple concurrent WhisperX workers.
  • WhisperX diarization depends on a Hugging Face token and will fall back cleanly if diarization is not configured.
  • Files copied into Windows-backed bind mounts may not always emit reliable filesystem events into Docker; EchoNotes now queues files already present at startup, but Linux-side writes remain the most reliable path.

QA Notes

Validated against the current Testing/ corpus on the CUDA image:

  • Standard PDF, DOCX, TXT, FLAC, and WAV samples completed successfully.
  • Encrypted PDFs were skipped gracefully without crashing the worker.
  • Short audio files produced transcript, summary, Obsidian note, and vault copies correctly.

Known issues from QA:

  • Very large audio can still hit GPU memory limits depending on model choice and VRAM.
  • Very large text inputs can take long enough to block a single-worker queue under slow local LLMs.