Repository navigation
Semantic AI Polish and QC
Language / שפה: English | עברית
This document specifies the architectural design, data contracts, CLI interface, and production implementation of the Semantic AI Polish & Subtitle QC Engine in RightSub (shipped in v1.3.0): the rightsub polish command.
The goal of this engine is to proofread, polish, and modernize existing Hebrew subtitles (legacy releases, imperfect human translations, or raw machine drafts — such as community scans like Star Wars 4K77) without re-translating the entire movie from scratch, correcting canon drift, gender agreement errors, and literal idioms while strictly preserving quality human phrasing (85%–90% preservation target).
Tens of thousands of community-uploaded Hebrew subtitle files exist for classic films and series. In 85%–90% of cases, the dialogue timing is accurate and basic phrasing is serviceable, yet they suffer from three pervasive semantic defects:
- Franchise & Canon Terminology Drift: Inconsistent or obsolete translations of lore terms (e.g. translating "Lightsaber" as "Laser Sword" / "חרב לייזר", or omitting military ranks like "Grand Moff").
- Grammatical Gender Flipping: Second-person English dialogue ("You") translated into masculine Hebrew verbs even when addressing female characters (e.g. Han or Luke addressing Princess Leia with masculine verbs).
- Anachronistic Registers & Literal Idiom Translation: Idiomatic expressions translated word-for-word, resulting in unnatural dialogue (e.g. "I have a bad feeling about this" translated literally instead of natural Hebrew cinematic prose).
Re-translating the entire film with an LLM burns unnecessary tokens, risks hallucinating or discarding authentic human dialogue choices, and can introduce timing drift. The Polish Engine operates on the principle of Minimal Edit Distance.
| Command | Primary Role | Modifies Text? | Uses AI? | Token Cost |
|---|---|---|---|---|
rightsub fix-plex |
Algorithmic BiDi (RLM), CP1255 recovery, and ad cleaning. | ❌ No | ❌ No | 0 |
rightsub qa |
Deterministic verification of line counts, timing, and homoglyphs. | ❌ Audit report only | ❌ No | 0 |
rightsub auto |
Autonomous runner (extracts, fixes Plex, or translates from scratch). | ✔️ Full (if translating) | ✔️ Yes (if needed) | 100% |
rightsub polish |
Proofreading, semantic gender correction & canon harmonization. | ✔️ Selective fixes only | ✔️ Yes (QC Reviewer) | ~15%–20% |
# Review and polish an existing Hebrew subtitle against master English:
rightsub polish "Star Wars (1977).he.srt" --en "Star Wars (1977).en.srt"
# Explicitly provide a Translation Bible (or auto-detected translation_bible.json):
rightsub polish "Star Wars (1977).he.srt" --bible translation_bible.json
# Query TMDb by specific ID for cast rosters and franchise canon:
rightsub polish "Gladiator.he.srt" --en "Gladiator.en.srt" --tmdb-id 98
# Cloud AI proofreading with automated Model Cascade (gemini-3.8-flash -> gemini-3.5-flash):
rightsub polish "Movie.he.srt" --gemini
# Run locally and 100% offline via Ollama:
rightsub polish "Movie.he.srt" --en "Movie.en.srt" --ollama --model qwen2.5:7b
# Dry run / Audit mode: generate a change report without modifying the subtitle file:
rightsub polish "Movie.he.srt" --en "Movie.en.srt" --diff-only[Movie.en.srt] ──┐
├──> [1. Bilingual Cue Alignment] ──> [2. Translation Bible & TMDb Ingestion]
[Movie.he.srt] ──┘ │
▼
[Master: Movie.he.polished.srt] <── [4. RLM & BiDi Enforcement] <─── [3. Constrained AI Polish Prompt]
│
└──> [5. Markdown Diff Report: Movie_polish_diff.md]
Pairs each English cue with its corresponding Hebrew cue based on timestamp intervals and sequence indexes:
{
"index": 142,
"timing": "00:18:22,100 --> 00:18:24,800",
"en": "Your father's lightsaber. This is the weapon of a Jedi Knight.",
"current_he": "חרב הלייזר של אביך. זה הנשק של אביר ג'דיי."
}Ingests companion translation_bible.json (or via -b/--bible) and queries TMDb API for the film's synopsis, genre categorization, and verified character roster with grammatical gender mapping (e.g., Luke Skywalker: male/אתה, Leia Organa: female/את). Also injects domain-specific glossaries for recognized franchises (Star Wars, Lord of the Rings, Harry Potter, Marvel Cinematic Universe).
Instructs the LLM to act strictly as a professional subtitling editor:
- The Conservation Rule: If the current Hebrew line is accurate, natural, and grammatically correct — do not touch it.
- Canon Correction: Align character names and established franchise terms to accepted modern canon.
- Gender Agreement: Align second- and third-person pronouns with the confirmed speaker and addressee.
- Flow & Syntax Modernization: Untangle awkward literal translations into natural cinematic dialogue.
- Subtitle Constraints: Strict limit of 38–40 characters per line, maximum 2 lines per block.
The LLM returns only the cues that required modification:
{
"cues": [
{
"index": 142,
"original_he": "חרב הלייזר של אביך. זה הנשק של אביר ג'דיי.",
"polished_he": "חרב האור של אביך. זהו נשקו של אביר ג'דיי.",
"reason": "Canon terminology (Lightsaber = חרב אור) + improved syntax"
}
]
}Produces a readable Markdown change log alongside the mastered subtitle:
# 📋 RightSub Polish Audit Report — Star Wars (1977)
- **Total Cues Audited:** 1,248
- **Preserved Intact:** 1,114 (89.3%)
- **Polished / Corrected:** 134 (10.7%)
### Line Modification Details:
| Cue | Original English | Previous Hebrew | Polished Hebrew | Reason |
| :---: | :--- | :--- | :--- | :--- |
| **#142** | "Your father's lightsaber." | "חרב הלייזר של אביך." | "חרב האור של אביך." | Canon terminology (Lightsaber) |
| **#315** | "Can you hear me, Princess?" | "אתה שומע אותי, נסיכה?" | "את שומעת אותי, נסיכה?" | Gender correction (Princess Leia is female) |
| **#520** | "I have a bad feeling about this." | "אני יש לי הרגשה רעה לגבי זה." | "יש לי תחושה רעה בקשר לזה." | Fixed literal clumsy machine syntax |RightSub operates an offline and online 3-tier classifier (domain_knowledge.py):
- Tier 1 (TMDb / Bible): Verified genre lists and character cast sheets directly from TMDb or local bibles.
- Tier 2 (Franchise Patterns): Regex and title matches across major film franchises (Star Wars, Marvel, Star Trek, Lord of the Rings, Harry Potter).
- Tier 3 (Offline Fingerprinting): Lexical keyword scans detecting sci-fi, military, medical, or legal dialogue inside the subtitle itself when offline.
To guarantee that cross-genre terms never corrupt unintended lines (e.g. translating "motor" as "motivator" outside of droid contexts), deterministic franchise rules require bidirectional validation:
- The Hebrew trigger must match the pattern.
- The master English cue must independently confirm the presence of the anchor phrase.
When dialogue spans multiple split cues (e.g. a sentence begun in cue #144 and finished in cue #145), standard LLMs often collapse the sentence into the first cue and translate upcoming dialogue lines into subsequent cues, causing a cascading off-by-one desynchronization.
-
Split-Cue Annotations: Cues are annotated with
split_partmarkers ("1/2","2/2"). - Drift Guard: Polished outputs are checked in real-time. If an edit leaks or quotes words from upcoming English master cues, the edit is automatically rejected and the authentic human translation is preserved.
When writing mastered subtitles or diff reports across SMB network mounts (e.g. Synology NAS, TrueNAS, QNAP), macOS SMB clients (smbfs) often fail to flush creation metadata, falling back to Apple's CoreFoundation epoch (2001-01-01).
RightSub explicitly calls os.utime() after all file write and backup operations, issuing an immediate kernel SYS_utimes call and SMB2 SET_INFO (FileBasicInformation) packet with the current system clock.
- Original subtitles (
Movie.he.srt) are safely backed up toMovie.he.original.srtbefore any replacement occurs. - Polished output files are passed through the SubRefine engine, guaranteeing UTF-8 encoding and idempotent RLM (
\u200F) injection for Plex, Infuse, and Apple TV playback. - Subtitle timecodes and cue indexes are strictly immutable.
RightSub Wiki — Subtitles Done Right. Powered by the SubRefine Algorithmic Engine & SubSwarm Multi-Agent AI.
- Home
- 📦 Global Installation Guide
- 🔰 Quickstart for Beginners
- BiDi & Plex Guide
- Pipeline Workflow
- On-Device STT & Sync (quicksubs)
- TMDb Metadata & Entity Resolution
- AI Assistants & Integration
- ⚖️ RightSub vs. Bazarr Comparison
- 🔄 Home Media & Download Integrations
- 🔮 Interactive CLI Specification
- 🔮 Setup & Health Wizard Specification
- 🔌 MCP Server Specification
- 💎 Semantic AI Polish & QC Specification
- 🐳 Docker Webhook Server Specification
- 🗺️ Product Roadmap
- 🚀 What's New & Release Notes
- Boston Legal Case Study
- CLI Reference
- דף הבית (Home HE)
- 📦 מדריך התקנה גלובלית והפצה
- 🔰 מדריך פשוט למתחילים
- מדריך BiDi ו-Plex
- תהליך עבודה מלא
- תמלול וסנכרון מקומי (quicksubs)
- אינטגרציית TMDb (עלילה ומגדר)
- חיבור לכלי בינה מלאכותית (AI)
- ⚖️ השוואה טכנית מול Bazarr
- 🔄 מדריך אינטגרציות ואוטומציה לשרתי מדיה
- 🔮 מפרט אשף פקודה אינטראקטיבי
- 🔮 מפרט אשף התקנה ואבחון
- 🔌 מפרט שרת MCP
- 💎 מפרט מנוע ליטוש סמנטי ו-QC
- 🐳 מפרט שרת Webhook וקונטיינר
- 🗺️ מפת דרכים ומעקב אבני-דרך
- 🚀 מה חדש ועדכוני גרסאות
- מקרה בוחן - בוסטון ליגל
- מדריך פקודות CLI