Skip to content

Semantic AI Polish and QC

omerninyo edited this page Sep 29, 2026 · 3 revisions

💎 Architectural Specification: Semantic AI Polish & Subtitle QC Engine (rightsub polish)

Language / שפה: English | עברית

This document specifies the architectural design, data contracts, CLI interface, and production implementation of the Semantic AI Polish & Subtitle QC Engine in RightSub (shipped in v1.3.0): the rightsub polish command.

The goal of this engine is to proofread, polish, and modernize existing Hebrew subtitles (legacy releases, imperfect human translations, or raw machine drafts — such as community scans like Star Wars 4K77) without re-translating the entire movie from scratch, correcting canon drift, gender agreement errors, and literal idioms while strictly preserving quality human phrasing (85%–90% preservation target).


1. Problem Statement & Engineering Motivation

Tens of thousands of community-uploaded Hebrew subtitle files exist for classic films and series. In 85%–90% of cases, the dialogue timing is accurate and basic phrasing is serviceable, yet they suffer from three pervasive semantic defects:

  1. Franchise & Canon Terminology Drift: Inconsistent or obsolete translations of lore terms (e.g. translating "Lightsaber" as "Laser Sword" / "חרב לייזר", or omitting military ranks like "Grand Moff").
  2. Grammatical Gender Flipping: Second-person English dialogue ("You") translated into masculine Hebrew verbs even when addressing female characters (e.g. Han or Luke addressing Princess Leia with masculine verbs).
  3. Anachronistic Registers & Literal Idiom Translation: Idiomatic expressions translated word-for-word, resulting in unnatural dialogue (e.g. "I have a bad feeling about this" translated literally instead of natural Hebrew cinematic prose).

Re-translating the entire film with an LLM burns unnecessary tokens, risks hallucinating or discarding authentic human dialogue choices, and can introduce timing drift. The Polish Engine operates on the principle of Minimal Edit Distance.


2. RightSub Subtitle Operation Comparison

Command Primary Role Modifies Text? Uses AI? Token Cost
rightsub fix-plex Algorithmic BiDi (RLM), CP1255 recovery, and ad cleaning. ❌ No ❌ No 0
rightsub qa Deterministic verification of line counts, timing, and homoglyphs. ❌ Audit report only ❌ No 0
rightsub auto Autonomous runner (extracts, fixes Plex, or translates from scratch). ✔️ Full (if translating) ✔️ Yes (if needed) 100%
rightsub polish Proofreading, semantic gender correction & canon harmonization. ✔️ Selective fixes only ✔️ Yes (QC Reviewer) ~15%–20%

3. CLI Command Interface

# Review and polish an existing Hebrew subtitle against master English:
rightsub polish "Star Wars (1977).he.srt" --en "Star Wars (1977).en.srt"

# Explicitly provide a Translation Bible (or auto-detected translation_bible.json):
rightsub polish "Star Wars (1977).he.srt" --bible translation_bible.json

# Query TMDb by specific ID for cast rosters and franchise canon:
rightsub polish "Gladiator.he.srt" --en "Gladiator.en.srt" --tmdb-id 98

# Cloud AI proofreading with automated Model Cascade (gemini-3.8-flash -> gemini-3.5-flash):
rightsub polish "Movie.he.srt" --gemini

# Run locally and 100% offline via Ollama:
rightsub polish "Movie.he.srt" --en "Movie.en.srt" --ollama --model qwen2.5:7b

# Dry run / Audit mode: generate a change report without modifying the subtitle file:
rightsub polish "Movie.he.srt" --en "Movie.en.srt" --diff-only

4. Pipeline Architecture

[Movie.en.srt] ──┐
                 ├──> [1. Bilingual Cue Alignment] ──> [2. Translation Bible & TMDb Ingestion]
[Movie.he.srt] ──┘                                                    │
                                                                      ▼
[Master: Movie.he.polished.srt] <── [4. RLM & BiDi Enforcement] <─── [3. Constrained AI Polish Prompt]
                 │
                 └──> [5. Markdown Diff Report: Movie_polish_diff.md]

Stage 1: Bilingual Cue Alignment

Pairs each English cue with its corresponding Hebrew cue based on timestamp intervals and sequence indexes:

{
  "index": 142,
  "timing": "00:18:22,100 --> 00:18:24,800",
  "en": "Your father's lightsaber. This is the weapon of a Jedi Knight.",
  "current_he": "חרב הלייזר של אביך. זה הנשק של אביר ג'דיי."
}

Stage 2: Translation Bible, Plot Synopsis & TMDb Ingestion

Ingests companion translation_bible.json (or via -b/--bible) and queries TMDb API for the film's synopsis, genre categorization, and verified character roster with grammatical gender mapping (e.g., Luke Skywalker: male/אתה, Leia Organa: female/את). Also injects domain-specific glossaries for recognized franchises (Star Wars, Lord of the Rings, Harry Potter, Marvel Cinematic Universe).

Stage 3: Constrained Polish Prompt Contract

Instructs the LLM to act strictly as a professional subtitling editor:

  1. The Conservation Rule: If the current Hebrew line is accurate, natural, and grammatically correct — do not touch it.
  2. Canon Correction: Align character names and established franchise terms to accepted modern canon.
  3. Gender Agreement: Align second- and third-person pronouns with the confirmed speaker and addressee.
  4. Flow & Syntax Modernization: Untangle awkward literal translations into natural cinematic dialogue.
  5. Subtitle Constraints: Strict limit of 38–40 characters per line, maximum 2 lines per block.

Stage 4: Structured AI Response Schema

The LLM returns only the cues that required modification:

{
  "cues": [
    {
      "index": 142,
      "original_he": "חרב הלייזר של אביך. זה הנשק של אביר ג'דיי.",
      "polished_he": "חרב האור של אביך. זהו נשקו של אביר ג'דיי.",
      "reason": "Canon terminology (Lightsaber = חרב אור) + improved syntax"
    }
  ]
}

Stage 5: Markdown Diff Audit Report

Produces a readable Markdown change log alongside the mastered subtitle:

# 📋 RightSub Polish Audit Report — Star Wars (1977)

- **Total Cues Audited:** 1,248
- **Preserved Intact:** 1,114 (89.3%)
- **Polished / Corrected:** 134 (10.7%)

### Line Modification Details:
| Cue | Original English | Previous Hebrew | Polished Hebrew | Reason |
| :---: | :--- | :--- | :--- | :--- |
| **#142** | "Your father's lightsaber." | "חרב הלייזר של אביך." | "חרב האור של אביך." | Canon terminology (Lightsaber) |
| **#315** | "Can you hear me, Princess?" | "אתה שומע אותי, נסיכה?" | "את שומעת אותי, נסיכה?" | Gender correction (Princess Leia is female) |
| **#520** | "I have a bad feeling about this." | "אני יש לי הרגשה רעה לגבי זה." | "יש לי תחושה רעה בקשר לזה." | Fixed literal clumsy machine syntax |

Stage 6: Hierarchical Domain Knowledge Engine

RightSub operates an offline and online 3-tier classifier (domain_knowledge.py):

  1. Tier 1 (TMDb / Bible): Verified genre lists and character cast sheets directly from TMDb or local bibles.
  2. Tier 2 (Franchise Patterns): Regex and title matches across major film franchises (Star Wars, Marvel, Star Trek, Lord of the Rings, Harry Potter).
  3. Tier 3 (Offline Fingerprinting): Lexical keyword scans detecting sci-fi, military, medical, or legal dialogue inside the subtitle itself when offline.

Bilingual Anchor Validation:

To guarantee that cross-genre terms never corrupt unintended lines (e.g. translating "motor" as "motivator" outside of droid contexts), deterministic franchise rules require bidirectional validation:

  • The Hebrew trigger must match the pattern.
  • The master English cue must independently confirm the presence of the anchor phrase.

Stage 7: Forward Anticipation Drift Guard & Split-Cue Partitioning

When dialogue spans multiple split cues (e.g. a sentence begun in cue #144 and finished in cue #145), standard LLMs often collapse the sentence into the first cue and translate upcoming dialogue lines into subsequent cues, causing a cascading off-by-one desynchronization.

  • Split-Cue Annotations: Cues are annotated with split_part markers ("1/2", "2/2").
  • Drift Guard: Polished outputs are checked in real-time. If an edit leaks or quotes words from upcoming English master cues, the edit is automatically rejected and the authentic human translation is preserved.

Stage 8: Network Storage (SMB/NAS) Metadata Synchronization

When writing mastered subtitles or diff reports across SMB network mounts (e.g. Synology NAS, TrueNAS, QNAP), macOS SMB clients (smbfs) often fail to flush creation metadata, falling back to Apple's CoreFoundation epoch (2001-01-01). RightSub explicitly calls os.utime() after all file write and backup operations, issuing an immediate kernel SYS_utimes call and SMB2 SET_INFO (FileBasicInformation) packet with the current system clock.


5. Non-Destructive Safety Guarantees

  • Original subtitles (Movie.he.srt) are safely backed up to Movie.he.original.srt before any replacement occurs.
  • Polished output files are passed through the SubRefine engine, guaranteeing UTF-8 encoding and idempotent RLM (\u200F) injection for Plex, Infuse, and Apple TV playback.
  • Subtitle timecodes and cue indexes are strictly immutable.

🌐 Navigation / ניווט

🇺🇸 English Documentation


🇮🇱 תיעוד בעברית


📦 Quick Links

Clone this wiki locally