Skip to content

Rage997/AutoSubs

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

6 Commits
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

autosubs

Automatic subtitle (.srt) generator CLI. Point it at video/audio files (or directories) and it writes a sidecar .srt subtitle file per input using a local faster-whisper speech-to-text model. Runs fully offline after the first model download.

Features:

  • Automatic spoken-language detection.
  • Translate-to-English mode (--translate).
  • Batch / recursive directory input.
  • Cross-platform: CPU int8 on macOS (Apple Silicon), CUDA on Linux/NVIDIA (untested but should work).

Try it without installing

With uv, you can run it once without installing anything permanently:

uvx --from autosubs-whisper autosubs VIDEO [VIDEO ...]

uvx fetches the package into a temporary environment, runs the autosubs command, and leaves nothing behind. (The --from flag is needed because the package is named autosubs-whisper (autosubs was already taken) while the command is autosubs.)

Install

uv tool install autosubs-whisper

This puts the autosubs command on your PATH (works on macOS, Linux, and Windows). pipx install autosubs-whisper works too.

For local development from a clone, use uv sync and uv run autosubs.

Usage

autosubs VIDEO [VIDEO ...]

The first run downloads the chosen Whisper model from the Hugging Face Hub (needs internet once); subsequent runs are offline.

Examples:

# Transcribe one file -> creates video.srt beside it
autosubs talk.mp4

# Recurse a directory, English translation, larger model
autosubs ./season1/ --translate --model large-v3

# Force source language, write into a separate folder, regenerate existing
autosubs clip.mkv --language ja --output-dir ./subs --overwrite

Options

Flag Default Description
--model small Whisper size/name: tiny, base, small, medium, large-v3, large-v3-turbo, distil-large-v3. Use large-v3 when a CUDA GPU is available.
--translate off Translate speech to English (Whisper task=translate).
--language auto Force source language ISO code (en, ja, ...). Omit to auto-detect.
--device auto auto (CUDA if available, else CPU), cpu, or cuda.
--compute-type int8 CTranslate2 compute type (int8, int8_float16, float16).
--beam-size 5 Beam search width.
--output-dir beside input Directory for .srt files.
--overwrite off Regenerate .srt files that already exist (default: skip).

By default, inputs whose target .srt already exists are skipped; pass --overwrite to regenerate.

Notes

  • Audio is decoded directly from video containers via PyAV (bundled with faster-whisper), so system FFmpeg is not required. If installed, ffmpeg is used as a fallback for exotic codecs PyAV cannot open.
  • On macOS only the CPU backend is available (no Metal GPU support in CTranslate2); int8 keeps it reasonably fast.
  • Accuracy vs. speed: Voice Activity Detection (VAD) is disabled by default to ensure maximum accuracy. All dialogue is transcribed, including quiet speech masked by music or sound effects. This is slow: expect approximately 1x realtime or slower (a 42-minute episode may take 40-60 minutes to transcribe). This prioritizes completeness over speed. If you need faster processing and can accept potentially missing very quiet dialogue, you can enable VAD by editing src/autosubs/transcribe.py to set vad_filter=True.

About

Whisper powered automatic movie subtitle generator

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages