-
Notifications
You must be signed in to change notification settings - Fork 0
Bazarr Comparison
Language / שפה: English | עברית
This technical audit is based on Bazarr's official documentation (Bazarr Subtitles Wiki) and a line-by-line inspection of Bazarr's open-source GitHub codebase (morpheus65535/bazarr).
Short Answer: Because Bazarr is an Ingestion Web Crawler & Downloader, while RightSub is a Semitic Linguistics, BiDi Engineering, and Mastering Factory.
Users frequently observe buttons in Bazarr's UI labeled Reverse RTL, Remove HI Tags, OCR Fixes, Translate, and Two-Point Fit, and assume Bazarr provides a complete subtitle solution for Hebrew. Inspecting the underlying source code reveals that these tools were inherited from the legacy Sub-Zero library, engineered specifically for Western European Latin languages. When executed on Semitic (Hebrew/Arabic) subtitles, they fail across four fundamental dimensions:
-
Physical BiDi Corruption: Bazarr's
Reverse RTLperforms a crude physical character swap instead of injecting logical Unicode directionality marks, corrupting subtitles on modern players (Apple TV, Infuse, Plex, Android TV). -
Total Blindness to Hebrew in Sanitization: Bazarr's OCR and Hearing Impaired filters contain zero Hebrew Unicode codepoints (
\u0590-\u05FF) in their regular expressions and dictionaries, silently skipping Hebrew files. - Flat, Gender-Blind Machine Translation: Translation tools execute a static 3-line prompt on massive 300-cue batches with zero pre-flight analysis, zero TMDb cast gender resolution, and zero context overlap, causing severe second-person grammatical gender flips ("אתה" vs. "את") and severed sentences.
- Torrent Seed Invalidation: In-place file edits inside media folders alter file sizes and cryptographic hashes, causing active torrent clients (qBittorrent, Transmission) to fail with I/O hash check errors.
Engineering Verdict: RightSub and Bazarr are not rivals; they are complementary. The optimal automated home media setup uses Bazarr to crawl and download community files from the web, and triggers RightSub as a Post-Processing hook to master, sanitize, and format the subtitles to broadcast standards.
-
In Bazarr Source Code (
custom_libs/subzero/modification/mods/common.py, lines 119–134):class ReverseRTL(SubtitleModification): identifier = "reverse_rtl" processors = [ NReProcessor(re.compile(r"(?u)(^([\s.!?:,'-]*)(.+?)(\s*)(-?\s*)$)"), r"\5\4\3\2", name="CM_RTL_reverse") ]
-
The Engineering Flaw:
- The routine assumes the target playback engine lacks Unicode BiDi support and performs a crude physical swap of line characters (
\5\4\3\2). - Modern media players (Plex, Apple TV, Infuse, Android TV, VLC) implement native Unicode Bidirectional Algorithm (UAX #9).
- When Bazarr's physical swap is applied, dialogue dashes jump from the start of the line to the end (
שלום-instead of- שלום). Furthermore, the regex trailing capture only matches hyphens (-?\s*$), completely ignoring trailing question marks (?) and exclamation points.
- The routine assumes the target playback engine lacks Unicode BiDi support and performs a crude physical swap of line characters (
-
The RightSub Solution (
SubRefine):- Keeps 100% of the underlying text in valid, logical UTF-8 sequence. Never physically reverses characters.
- Injects invisible Unicode Right-to-Left Marks (
\u200F- RLM) adjacent to neutral punctuation and dialogue dashes in strict compliance with Unicode UAX #9. - Result: Flawless rendering on Apple TV, Plex, Infuse, and Smart TVs without character inversions.
-
In Bazarr Source Code (
custom_libs/subzero/modification/mods/hearing_impaired.py, lines 60–64):NReProcessor(re.compile(r'(?sux)-?%(t)s[\"\']*\[(?=[^\[\]]{3,})[A-Za-zÀ-ž0-9\s\'\".:-_&+]+[\)\]][\"\'*[\s:]*%(t)s' % {"t": TAG}), "", name="HI_brackets")
-
The Engineering Flaw:
- The bracket regex character class
[A-Za-zÀ-ž0-9...]strictly matches Latin and Western European characters. -
Hebrew Unicode codepoints (
\u0590-\u05FF) are completely excluded. Tags such as[מוזיקה מתנגנת],(טלפון מצלצל), or[צעקות]are never matched and remain intact. - Additionally, the All-Caps sound description list is hardcoded to exactly 19 English-only words (
LAUGH,SCREAM,DOOR, etc.).
- The bracket regex character class
-
The RightSub Solution (
junk_and_sdh_cleaning.py):- Full native support for Hebrew and English auditory patterns across all bracket types (
[],(),{}). - Removes musical notation (
♪) while strictly preserving cue timing intervals. - Whitelist protection (
test_legitimate_dialogue_not_flagged_as_ad) to ensure short dialogue lines are never mistakenly stripped.
- Full native support for Hebrew and English auditory patterns across all bracket types (
-
In Bazarr Source Code (
custom_libs/subzero/modification/mods/ocr_fixes.py):data_dict = OCR_fix_data.get(parent.language.alpha3t) if not data_dict: logger.debug("No SnR-data available for language %s", parent.language) return
-
The Engineering Flaw:
- Inspecting
dictionaries/data.pyreveals dictionary tables for only 16 European languages:
['bos', 'dan', 'deu', 'eng', 'fin', 'fra', 'hrv', 'hun', 'nld', 'nob', 'nor', 'por', 'rus', 'spa', 'srp', 'swe']. -
Hebrew (
heb) is totally unsupported. When clicked on a Hebrew subtitle, Bazarr outputs a debug log entry and silently exits without modifying anything.
- Inspecting
-
The RightSub Solution (v1.3.0):
-
Universal Homoglyph Normalization: Automatically identifies and replaces foreign character intrusions introduced by legacy bitmap OCR (e.g. Cyrillic
сresembling Latincor Hebrew characters). -
Hebrew Typographical Normalization: Converts ASCII double quotes into canonical Hebrew Gershayim (
״/\u05F4) for acronyms (e.g., עו״ד, ארה״ב, ד״ר). -
Semantic AI Polish (
rightsub polish): Uses LLMs and 3-tier domain lore (domain_knowledge.py) to audit spelling, slang, and franchise canon.
-
Universal Homoglyph Normalization: Automatically identifies and replaces foreign character intrusions introduced by legacy bitmap OCR (e.g. Cyrillic
-
In Bazarr:
- Bazarr crawls files from community websites (OpenSubtitles, Subscene, Torec) and saves them with their original encoding.
- Many older Israeli subtitle archives were saved in Windows-1255 (CP1255) or ISO-8859-8.
- When modern streaming players read CP1255 files under UTF-8 assumptions, subtitles display as unreadable gibberish (
àðé øåöäinstead ofאני רוצה). Bazarr has no automated Semitic encoding detection or conversion pipeline.
-
The RightSub Solution:
- The Auto-Charset engine evaluates all input files with statistical heuristic confidence, automatically detecting CP1255 and ISO-8859-8 encodings and converting them to clean, modern UTF-8.
-
In Bazarr Source Code (
custom_libs/subzero/modification/mods/two_point_fit.py):parent.f.shift(...) parent.f.transform_framerate(float(kwargs.get("from")), float(kwargs.get("to"))) parent.f.shift(...)
-
The Operational Difference:
- The underlying linear time scaling and offset math matches RightSub's
timing_and_sync.py. - The Bazarr Bottleneck: Requires manual user interaction: opening the video in an external player, seeking out two spoken lines, recording 4 separate timestamps, typing them into a web GUI modal, and clicking save.
- The underlying linear time scaling and offset math matches RightSub's
-
The RightSub Solution:
-
Single-Line Headless CLI: FPS conversions and millisecond offsets execute in a single headless command (
rightsub sync). -
Hardware-Accelerated Alignment (
quicksubs): Performs local on-device speech-to-text audio alignment against the video's audio track via Apple Silicon Neural Engine/Metal or local Whisper models — zero manual timestamp hunting and zero cloud API expenses.
-
Single-Line Headless CLI: FPS conversions and millisecond offsets execute in a single headless command (
This is the primary architectural differentiator refuting the misconception that Bazarr's "Translate" button makes an external tool redundant.
-
Google Translator (
google_translator.py): Dispatches individual cues in parallel viaThreadPoolExecutor(max_workers=10). Every line is translated in complete isolation: zero conversational context, total grammatical gender blindness, and broken sentence flow. -
Gemini Translator (
gemini_translator.py, lines 135–150): Sends monolithic 300-cue blocks using a static, 3-line prompt:You are an assistant that translates subtitles to {language}. Dialogs must be translated as they are without any changes. If a line has a comma or multiple sentences, try to keep one line to about 40-50 characters.- Zero Pre-Flight Intelligence: No TMDb character cast query, no speaker extraction, and no awareness of speaker identities or genders.
- Zero Context Overlap: If a dialogue sentence is split across subtitle batches, the first half is translated in batch A and the second half in batch B without conversational context.
-
Destructive BiDi Encodings: Bazarr wraps Hebrew lines in legacy Unicode 2.0
\u202b...\u202c(RLE/PDF) directional embedding codes. In modern media players (Apple TV, Infuse, Plex), these obsolete marks physically invert numbers, times, and currency symbols.
-
Phase 1: Pre-Flight Bible & TMDb Entity Resolution (
03_generate_bible.py):- Before submitting a single token to the LLM, RightSub scans the English SRT for explicit speaker tags (
ALAN:,[DENNY]), character names, and recurring honorifics (Judge,Counselor,Dr.). - Executes an automated pre-flight query against the TMDb API to resolve the verified cast roster, character names, and official genders (Male/Female).
- Generates and locks
translation_bible.json, guaranteeing that character gender and nomenclature remain 100% consistent across entire 24-episode seasons.
- Before submitting a single token to the LLM, RightSub scans the English SRT for explicit speaker tags (
-
Phase 2: Master System Prompt & Context Overlap (
09_prompt_builder.py):- Explicit Vocative Agreement Rules: Enforces strict Semitic 2nd-person gender inflection ("אתה" vs. "את", "אתה רוצה" vs. "את רוצה") based on speaker and addressee continuity.
-
SDH Filtered as Context-Only: Speaker tags like
[Tara]orALAN:are supplied as context to inform the LLM of conversational participants, while the prompt strictly forbids outputting these metadata tags into final subtitles. - Sliding Context Overlap: Every batch receives 3–5 overlapping dialogue lines from preceding and succeeding cues, preventing fragmented sentences.
-
Domain Knowledge & Canon Lore (
domain_knowledge.py): Dynamically injects genre glossaries (legal courtroom terms, medical jargon, sci-fi franchises).
-
Phase 3: Semantic AI Polish (
19_polish_and_qc.py):- If Bazarr already downloaded an existing community subtitle that suffers from machine-translation errors, RightSub does not need to re-translate the file from scratch.
- The
rightsub polishengine operates on Minimal Edit Distance, auditing the Hebrew subtitle against master English dialogue to preserve 85%–90% of valid human phrasing while surgically fixing gender flips and terminology at ~15% of the token cost.
-
In Bazarr:
- When Bazarr edits, converts, or synchronizes a subtitle file inside a media folder, it alters the file in-place.
- This mutates the file size and payload content. BitTorrent clients (qBittorrent, Transmission) detect hash mismatches during re-checks and immediately halt torrent seeding with I/O errors.
-
The RightSub Solution:
- Operates on a Strictly Non-Destructive Seed-Safe Model by default.
- Leaves the original downloaded subtitle 100% bit-for-bit intact, creating an external
<stem>.he.srtthat Plex and Infuse automatically recognize and prioritize.
-
Bazarr:
- Designed as an always-on server daemon or continuous Docker container running 24/7.
- Occupies port 6767, maintains background worker loops, and consumes continuous RAM (150–300 MB).
-
RightSub:
- Zero-Daemon Architecture.
- Lightweight CLI utility with zero background processes.
- Idle resource consumption: 0% CPU and 0 MB RAM. Wakes up on event triggers or manual calls, masters the subtitle in ~0.2s, and immediately terminates.
| Capability / Dimension | Bazarr | RightSub (v1.3.0) |
|---|---|---|
| Primary Pipeline Role | Automated web crawler & downloader for existing community subtitles (Ingest). | Mastering, Semitic linguistics, BiDi RLM, AI proofreading, and headless automation (Mastering). |
| Plex & Infuse BiDi Handling | ❌ Broken physical character swap (common.py) or legacy RLE tags (\u202b). |
✔️ SubRefine engine injecting invisible Unicode RLM (\u200F) marks (UAX #9). |
| Hearing Impaired / SDH Sanitization | ❌ Latin characters only; blind to Hebrew auditory descriptions. | ✔️ Native bilingual SDH sanitizer with dialogue protection. |
| OCR Typo Repair | ❌ Hebrew is completely missing from dictionary (data.py). |
✔️ Homoglyph normalization from 7 scripts + Hebrew Gershayim normalization. |
| Legacy Mojibake Repair (CP1255) | ✔️ Auto-Charset engine detecting Windows-1255 and converting to UTF-8. | |
| Timing Synchronization (Sync) | ✔️ Direct CLI flags (rightsub sync) + hardware-accelerated STT (quicksubs). |
|
| Pre-Flight Architecture & Bible | ❌ None. Static 4-line prompt with zero domain knowledge. | ✔️ TMDb cast resolution, speaker extraction, locked translation_bible.json. |
| Context Overlap Across Batches | ❌ Flat 300-line batches. Sentences severed without context. | ✔️ 3–5 cue sliding overlap windows guaranteeing sentence continuity. |
| Vocative & Gender Agreement | ❌ Completely blind (defaults to masculine, frequent gender flips). | ✔️ Explicit 2nd-person gender rules based on speaker continuity. |
| Subtitle Proofreading / Polish | ❌ None. Requires translating from scratch. | ✔️ rightsub polish with Minimal Edit Distance (85%–90% preservation). |
| Torrent Seeding Protection (Seed-Safe) | ❌ Modifies files in-place, corrupting torrent hashes. | ✔️ Non-destructive duplication to <stem>.he.srt (Bit-for-Bit intact). |
| NAS / SMB Metadata Synchronization | None (susceptible to macOS SMB 2001 epoch bug). | Explicit os.utime() metadata sync for immediate NAS indexing. |
| Automated 24/7 Web Crawling | ✔️ Industry standard (crawls 30+ providers linked to Sonarr/Radarr). | ❌ Does not crawl pirate sites; processes local files or generates subtitles via AI. |
| System Resource Overhead | 24/7 Docker container / background daemon (RAM/CPU). | Zero-daemon CLI running purely on event hooks (0% idle CPU, 0 MB RAM). |
The most reliable media server architecture pairs Bazarr for ingestion with RightSub for mastering:
In Bazarr (Settings -> Subtitles -> Post-processing -> Custom Post-Processing):
rightsub auto "{{subtitles_path}}"- Acquisition: Bazarr monitors media libraries and downloads raw community subtitles from web providers.
- Autonomous Mastering (0.2s): RightSub triggers immediately upon download, converts CP1255 to clean UTF-8, injects invisible RLM marks for Plex/Apple TV, removes SDH noise and ad spam, preserves torrent seeding, and updates NAS timestamps.
- AI Fallback & Polish: If Bazarr finds no subtitle on the web, or retrieves a low-grade machine translation, RightSub can translate or polish the media file directly at broadcast quality.
RightSub Wiki — Subtitles Done Right. Powered by the SubRefine Algorithmic Engine & SubSwarm Multi-Agent AI.
- Home
- 📦 Global Installation Guide
- 🔰 Quickstart for Beginners
- BiDi & Plex Guide
- Pipeline Workflow
- On-Device STT & Sync (quicksubs)
- TMDb Metadata & Entity Resolution
- AI Assistants & Integration
- ⚖️ RightSub vs. Bazarr Comparison
- 🔄 Home Media & Download Integrations
- 🔮 Interactive CLI Specification
- 🔮 Setup & Health Wizard Specification
- 🔌 MCP Server Specification
- 💎 Semantic AI Polish & QC Specification
- 🐳 Docker Webhook Server Specification
- 🗺️ Product Roadmap
- 🚀 What's New & Release Notes
- Boston Legal Case Study
- CLI Reference
- דף הבית (Home HE)
- 📦 מדריך התקנה גלובלית והפצה
- 🔰 מדריך פשוט למתחילים
- מדריך BiDi ו-Plex
- תהליך עבודה מלא
- תמלול וסנכרון מקומי (quicksubs)
- אינטגרציית TMDb (עלילה ומגדר)
- חיבור לכלי בינה מלאכותית (AI)
- ⚖️ השוואה טכנית מול Bazarr
- 🔄 מדריך אינטגרציות ואוטומציה לשרתי מדיה
- 🔮 מפרט אשף פקודה אינטראקטיבי
- 🔮 מפרט אשף התקנה ואבחון
- 🔌 מפרט שרת MCP
- 💎 מפרט מנוע ליטוש סמנטי ו-QC
- 🐳 מפרט שרת Webhook וקונטיינר
- 🗺️ מפת דרכים ומעקב אבני-דרך
- 🚀 מה חדש ועדכוני גרסאות
- מקרה בוחן - בוסטון ליגל
- מדריך פקודות CLI