Skip to content

Releases: stremovskyy/transcriptor

v1.1.0 - Enhanced Keyword Handling & Audio Preprocessing

Choose a tag to compare

@stremovskyy stremovskyy released this 08 Mar 07:21
e03a263

Release Notes - v1.1.0

Enhanced Keyword Handling & Audio Preprocessing 🚀

This release introduces powerful new keyword capabilities and audio preprocessing features, along with significant reliability improvements for production workloads.

✨ New Features

Keyword Spotting 2.0

  • ✅ Negated keyword support: Use !keyword syntax to detect segments without specific terms
  • 🎯 Improved accuracy: Better handling of homophones through fuzzy matching enhancements
  • 📊 Enhanced logging: Detailed keyword processing metrics in request logs

Audio Preprocessing Pipeline

  • 🎧 Smart audio enhancement: New preprocessing options including:
    • DC offset removal
    • Normalization & dynamic range compression
    • Noise reduction & silence trimming
    • Pre-emphasis filtering
  • ⚙️ Configurable workflow: Enable/disable steps via environment variables
  • 🖥️ UI integration: Toggle preprocessing in web interface

Gemma Improvements

  • 🔒 Safer model loading: Better error handling for GPU memory constraints
  • 📈 Performance metrics: Detailed memory usage tracking during initialization

🛠️ Improvements

  • 💾 GPU memory management: Optimized VRAM allocation logging and fallback logic
  • ✨ Type hint modernization: Better developer experience across codebase
  • 📄 Documentation overhaul: Updated README with preprocessing guides and new API examples
  • 🌐 UI enhancements: File preprocessing toggle and improved status displays

⚠️ Updated Known Issues

  • Gemma-2b-it still requires >8GB GPU RAM for optimal performance
  • AAC support continues to require FFmpeg preprocessing
  • Keyword confidence thresholds may need adjustment for technical vocabulary

🚨 Migration Note
Add these new environment variables to your .env file for preprocessing controls:

AUDIO_ENABLE_PREPROCESSING=true
AUDIO_ENABLE_DC_OFFSET=true
AUDIO_ENABLE_NORMALIZATION=true
AUDIO_PRE_EMPHASIS=0.97

v1.0.0 - Initial Release

Choose a tag to compare

@stremovskyy stremovskyy released this 04 Mar 09:05
168483d

Release Notes - v1.0.0 (Initial Release)

The First Stable Release 🎉
Production-Ready Audio Transcription & Text Reconstruction Service


🚀 New Features

Core Functionality

  • Whisper Model Integration: Full support for OpenAI's Whisper speech-to-text models (base, small, medium, large)
  • Gemma Reconstruction: Text refinement using Google's Gemma-2b-it LLM
  • Multi-Format Processing:
    • Direct file uploads via /transcribe endpoint
    • Remote URL processing via /pull endpoint

API Capabilities

  • 🔐 API key authentication middleware
  • 🚦 Rate limiting (10,000 requests/hour)
  • 🌐 Multi-language support (Ukrainian/Russian primary focus)
  • 🔍 Keyword spotting with confidence scoring
  • ⏱️ Processing time metrics in all responses

Infrastructure

  • 🧠 Smart model caching system with:
    • Automatic GPU/CPU fallback
    • Memory optimization
    • Concurrent request safety
  • 📊 Detailed logging (app.log & request logs)
  • 🐳 Production-ready Gunicorn configuration

⚠️ Known Issues

Performance

  • Initial model load time can be slow (~30-60s for large Whisper models)
  • Gemma models requires authorization for downloading
  • Gemma-2b-it requires >8GB GPU RAM for optimal performance
  • No native AAC audio support - requires FFmpeg preprocessing

Limitations

  • Maximum file size hard-capped at 50MB
  • Keyword spotting accuracy decreases with homophones
  • UI only supports basic upload functionality

🛠️ Upgrade Guide

New Requirements

# Required system packages
sudo apt-get install ffmpeg python3-dev

Critical Configuration

# .env changes from pre-1.0 versions
API_KEY_ENABLED=true # Now required by default
MAX_CONTENT_LENGTH=52428800 # Explicit size limit

📦 Installation

# For first-time users
git clone https://github.com/yourorg/audio-transcription-service.git
cd audio-transcription-service
pip install -r requirements.txt

🙌 Acknowledgments

  • OpenAI for Whisper speech recognition models
  • Google Research for Gemma language models
  • Flask & Torch communities for foundational libraries

First production release marks completion of core feature set. Subsequent releases will focus on performance optimization and expanded language support.