Skip to content

Latest commit

 

History

40 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

English · 中文

Index-Translate

A Multilingual Translation Model Family
Text, Speech, Controlled Dubbing, and Long-Document Translation

🌐 Demo · 🤗 Hugging Face ·  ModelScope · 📚 Report · 🤗 Papers · 🤗 Benchmarks · 🐧 QQ

Index-Translate is a family of multilingual translation models built on Qwen3.5. The text models cover 150 languages and follow translation instructions such as terminology, formatting, and content-preservation requirements. The family extends this foundation to speech, syllable-controlled translation, and full-document translation.

  • Index-Translate translates text, structured content, and community expressions.
  • Index-Echo produces translated subtitles or speech conditioned on the source speaker's voice.
  • Index-Homura adjusts translations toward a specified target syllable count.
  • Index-NativeLong translates complete documents with context across passages.

Seven-category comparison of Index-Translate 35B-A3B preview, 9B, and 2B

The radar includes 35B-A3B (preview), 9B, and 2B, with fixed per-axis min–max ranges across all 14 models. Its seven axes are WMT, FLORES, instruction following, low-resource translation, subtitles, MEME, and books/fiction. Instruction following averages instTrans and IFMTBench IFscore. The normalized scale is not an accuracy percentage. The gray dashed line combines the best non-Index score on each axis and does not represent one model. Raw category scores · Figure notes · Individual benchmark results.

News · Models · Quick start · Instruction Following · Examples · Evaluation · Benchmarks · Applications · TODO · Papers and citation

News

  • 2026-10-04: released free public API endpoints on index-translate.bilibili.com/v1 for Index-Translate-35B-A3B. Fully OpenAI-compatible. Try it with call_api.py.
  • 2026-10-04: released four Index-Translate benchmarks, with datasets/metadata on Hugging Face and evaluation scripts and guides on GitHub.
  • 2026-10-03: released official quantized builds across the family on Hugging Face and ModelScope — GGUF for llama.cpp local inference, alongside FP8 (W8A8) and NVFP4 (W4A4, Blackwell-optimized) for vLLM serving.
  • 2026-09-30: released Index-Translate, with 2B / 9B / 35B-A3B (preview) text-model weights on Hugging Face and ModelScope, the technical report, and the online demo.

TODO

  • Release official quantized builds (GGUF for llama.cpp local inference, FP8 and NVFP4 for vLLM serving) for the whole family.
  • Release the official version of Index-Translate-35B-A3B.
  • Release instTrans, MEME, SandGlass and NAtIveLong data/metadata and evaluation code; see Benchmarks.
  • Add support for more languages to Index-Echo.
  • Release larger models.

Models

The links below provide 2B, 9B, and 35B-A3B (preview) text-model checkpoints. Evaluation results are included in the comparison tables.

Model Task and released package coverage Hugging Face ModelScope Inference
Index-Translate Text translation and instructions across 150 languages 2B · 9B · 35B-A3B (preview) 2B · 9B · 35B-A3B (preview) Guide
Index-Echo S2TT Speech → subtitles; packaged script: zh→en/ja/es 2B · 9B 2B · 9B Guide
Index-Echo S2ST Speech → speech; zh→en/es/ja, en→zh/es/ja 2B · 9B 2B · 9B Guide
Index-Homura Translation with a target syllable count 2B · 9B 2B · 9B Guide
Index-NativeLong Long documents; fixed templates: zh↔en, zh↔ja 2B · 9B 2B · 9B Guide

Naming: Index-NativeLong is published under the model IDs IndexTeam/Index-Nailong-2B and IndexTeam/Index-Nailong-9B. Use those IDs in commands. Language support for the speech and long-document packages is listed separately from the text models' 150-language coverage.

Quantized builds: every model above is also published in GGUF (llama.cpp local inference; all bit-widths in one repository per model, vision mmproj included where applicable), FP8 (compressed-tensors W8A8, ready for vLLM serving), and FP4 (compressed-tensors NVFP4 W4A4, for Blackwell GPUs). For the Index-Echo speech models, the GGUF repositories contain the text LLM backbone only, while the FP8/FP4 repositories ship the complete pipeline with a quantized LLM. All repositories are mirrored on ModelScope.

Model GGUF (llama.cpp) FP8 (vLLM) FP4 (vLLM, Blackwell)
Index-Translate 2B · 9B · 35B-A3B (preview) 2B · 9B · 35B-A3B (preview) 2B · 9B · 35B-A3B (preview)
Index-Homura 2B · 9B 2B · 9B 2B · 9B
Index-NativeLong 2B · 9B 2B · 9B 2B · 9B
Index-Echo S2TT 2B · 9B (LLM backbone) 2B · 9B (full pipeline) 2B · 9B (full pipeline)
Index-Echo S2ST 2B · 9B (LLM backbone) 2B · 9B (full pipeline) 2B · 9B (full pipeline)

Inference

Quick start

Option 1: Free Online API (Zero GPU Setup)

You can call our free online API directly without local GPUs:

# Using the zero-dependency Python script
python inference/llm/call_api.py "你好,世界。今天天气不错,我们去公园散步吧。" --target en

# Or using curl (OpenAI-compatible)
curl https://index-translate.bilibili.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Index-Translate-35B-A3B",
    "messages": [{"role": "user", "content": "请将以下文本翻译为英语,直接输出翻译结果,不要进行任何解释。\n\n你好,世界。"}],
    "max_tokens": 100
  }'

Option 2: Self-hosted local vLLM

Start with the 2B text model on a CUDA GPU using a vLLM build with Qwen3.5 support. From a terminal:

git clone https://github.com/bilibili/Index-Translate.git
cd Index-Translate
pip install -U vllm
pip install -r inference/llm/requirements.txt
vllm serve IndexTeam/Index-Translate-2B --host 127.0.0.1 --port 8000 --max-model-len 4096

From the same repository directory in another terminal:

python inference/llm/translate.py \
  "你好,世界。今天天气不错,我们去公园散步吧。" \
  --target en --model IndexTeam/Index-Translate-2B

An output recorded with the released 2B model is:

Hello, world. The weather is nice today. Let's go for a walk in the park.

See captured cases, text inference, and prompt examples. The 4,096-token setting above is a short-text example; the serving presets are listed in the table below.

For audio, use the dedicated S2TT subtitle guide or S2ST dubbing guide.

Instruction Following & Constrained Translation

Index-Translate deeply integrates instruction-following capabilities (instTrans). It supports both hard constraints (format preservation, strict terminology glossary enforcement) and soft constraints (tone & style, domain disambiguation) out of the box via translate.py and syllable_translate.py:

# Hard constraint 1: Strict terminology glossary enforcement (-g / --glossary)
python inference/llm/translate.py \
  "王平仲采用了更加昂贵的碳纤维材料。碳纤维的好处就是它抗裂缝。" \
  --target en -g "碳纤维:carbon fiber, 抗裂缝:crack resistance"

# Hard constraint 2: Format & structure preservation (-H / --hard)
python inference/llm/translate.py \
  '{"user_id": 1024, "event": "purchase", "message": "您的订单已支付完成。"}' \
  --target en -H "保留源文中的 JSON 格式标记不变"

# Soft constraint 1: Tone and style adjustment (-S / --soft) + genre (-d / --genre)
python inference/llm/translate.py \
  "今天下午的会议临时取消了,改天我们再碰一下商量。" \
  --target en -d "商务邮件" -S "调整为严谨、正式、礼貌的商务公文风格"

# Soft constraint 2: Domain context and word-sense disambiguation (-S / --soft)
python inference/llm/translate.py \
  "The plant is operating at full capacity after the spring upgrade." \
  --target zh -d "工业制造" -S "语境为工业制造与重工厂房领域,准确消歧专有名词(如 plant 译为工厂而非植物)"

# Syllable control synergy: Index-Homura strictly respects syllable budgets while embedding glossaries
python inference/llm/syllable_translate.py \
  "我们今天去看电影吧" --syllables 7 --target en --glossary "电影:cinema"

instTrans Specification & Constraint Details:

  • Canonical Prompt: Automatically formatted by the client into the instTrans benchmark structure (【源文】 + numbered 1. 【硬性要求】... / 2. 【注意】... + suffix instructions).
  • Hard Constraints: Structural formatting (JSON/CSV/code/placeholders), terminology glossaries, social elements, and syllable ordering. Binary gated ($g_{\mathrm{hard}}$); any single failure zeroes the instance score.
  • Soft Constraints: Tone/style adaptation, contextual sense disambiguation, cross-sentence consistency, and LaTeX preservation. Evaluated on a graded scale ($q_{\mathrm{soft}}$).
  • See the text inference guide and instruction cases for full examples.

Default inference settings

This is the shared reference for the repository clients and the released Echo packages. Translate covers 2B / 9B / 35B-A3B (preview); the other families cover 2B / 9B. Decoding defaults are shared across sizes within each family.

Not set means the client inherits the backend/model configuration; — means the setting does not apply. Text-generation rows describe transcription/translation for Echo; speech generation has separate rows.

Setting Index-Translate Index-Homura Index-NativeLong Echo S2TT Echo S2ST
Entry point translate.py syllable_translate.py doc_translate.py s2tt.py → package infer.py dub.py → package DubbingBridgeModel
Default checkpoint 9B 9B 9B (Index-Nailong) 2B Local ./Index-Echo-S2ST-2B
Text decoding Greedy Sampling Greedy Greedy (do_sample=False) Greedy (do_sample=False)
temperature 0 0.3 0 0 0 (text)
top_p Not set Not set 1 Not set Not set (text)
top_k Not set Not set -1 Not set Not set (text)
min_p Not set Not set 0 Not set Not set (text)
presence_penalty Not set Not set 0 — —
repetition_penalty Not set Not set 1 Not set Not set (text)
seed Not set Not set 42 Not set 42 (speech)
Thinking enable_thinking=False enable_thinking=False enable_thinking=False Empty <think> block prefilled Empty <think> block prefilled
Text output budget max_tokens=1024 max_tokens=max(512, 3 * len(text)) max_tokens omitted; server selects the cap max_new_tokens=2000 per window max_new_tokens=1024
Text stop conditions Not set Not set stop_token_ids=[248044, 248046]; ignore_eos=False Tokenizer EOS / <|im_end|> Tokenizer EOS / <|im_end|>
Output streaming No No stream=True Sequential window results Speech stream=False
serve_vllm.sh context (input + output) 2B / 9B: 32768; no 35B preset 32768 2B: 262144; 9B: 229376 — —
Default language / constraint Source auto, target en Target en; --syllables required zh-en zh-en --lang required; source inferred as zh/en
Audio window / history — — — --max-win 60 seconds; --ctx-k 5 prior windows Utterances ≤30 seconds recommended; chunk=False
Speech sampler — — — — Fixed sampling=25; CosyVoice RAS defaults top_p=0.8, top_k=25
Speech-token budget — — — — Maximum min(1500, 20 * m); minimum 2 * m
Speech speed / sample rate — — — — speed=1.0; 24000 Hz
Main overrides --model, --instruction, --glossary, --temperature, --max-tokens --model, --syllables, --glossary, --temperature, --max-tokens --model, --direction, --max-tokens --size, --temperature, --max-new-tokens, --max-win, --ctx-k, --glossary --model-dir, --lang; lower-level API for budgets / seed
  • Text clients: default to http://127.0.0.1:8000/v1 with API key EMPTY. Use --base-url / --api-key or OPENAI_BASE_URL / OPENAI_API_KEY; --model overrides INDEX_MODEL and the default checkpoint. Serve 35B-A3B manually and select it with --model IndexTeam/Index-Translate-35B-A3B-preview.
  • Budgets: Homura's len(text) is the Python character count after trimming input. NativeLong's omitted max_tokens leaves the output cap to the server; context capacity and server limits still apply. Pass a positive --max-tokens for an explicit cap. Context includes the full prompt and generated output; override serving limits with --max-model-len.
  • Echo S2ST: m is the aligned target-text token count. sampling=25 is a fixed package-code argument, not a CLI option. The public DubbingBridgeModel.dub wrapper exposes neither seed nor token budgets; use the lower-level extract / synth API documented in the model card. chunk=True is not implemented.
  • Scope: these are usage defaults. The technical report's benchmark decoding and context settings are documented separately in Evaluation.

Examples

These examples are drawn from the official demo. Outputs below illustrate individual cases; full comparisons and task settings are available on the demo and in the report.

Translate text while preserving a hashtag

Task: translate this JSON into Korean, preserving its structure, stars, and the Chinese hashtag.

{"title": "⭐2月13日例行维护公告⭐", "content": "#热血航线大和登场#"}

Index-Translate-9B:

{"title": "⭐2월 13일 정기 점검 공지⭐", "content": "#热血航线大和登场#"}

The title is translated while the requested hashtag remains unchanged. Try text translation.

Keep the meaning of a community expression

Source: 狒瘾犯了就去打 — in this gaming context, “狒瘾” refers to the urge to play Final Fantasy XIV.

Model English output
Index-Translate-9B When the FFXIV itch hits, just go play.
Index-Translate-2B Go play when my FFXIV addiction kicks in.
Hy-MT2-7B If you get monkey addiction, go fight.

Set the translation's syllable budget

Source: 生活两天,是一种什么体验. Index-Homura-9B produces different wording for three requested lengths:

Target / observed syllables English output
10 / 10 to live for two days. What would that be like?
14 / 14 What would it be like to live there for two days, I wonder?
18 / 18 What would it be like to live there for two days, trying to get by somehow?

These three examples meet their targets; syllable control is approximate in general, and spoken duration also depends on delivery. Try Index-Homura.

Keep a character's name consistent across a document

In a roughly 32K-token fantasy document, “王妃” is a character's name. The following extracts compare native full-document translation with the same 9B model in a chunked workflow using neighboring context and an automatic glossary.

Source position Index-NativeLong-9B full document Same 9B with chunking
31.3% Mentor Wang Fei Instructor Wangfei
56.6% Wang Fei the Dean
84.8% Wang Fei The Queen Consort

Positions are measured by source characters. The full-document output keeps the name at these locations; the demo also shows how an externally supplied glossary can repair the chunked result. Try Index-NativeLong.

Watch speech translation

Index-Echo evaluation

The S2ST panels compare the deployed 2B system, a pipeline, and SeamlessM4T-v2. This demo comparison is separate from the six-direction matched study in the report; SeamlessM4T-v2 does not clone the source voice.

Speech-to-speech dubbingMultilingual subtitles
Watch the Index-Echo English-to-Japanese dubbing demo Watch the Index-Echo multilingual subtitle demo
English video with Japanese dubbingVideo with translations in multiple languages

Click either preview to watch the video, or open Index-Echo. The website demonstration and the packaged inference interfaces have different language coverage; the model table lists the released interfaces.

Evaluation

Updated text translation benchmarks

The charts reproduce the updated demo comparison. *35B-A3B is the preview model. The instruction panels average instTrans and IFMTBench: Quality combines instTrans quality and IFMTBench XCOMET-XXL, while IFscore averages their instruction scores. Individual benchmark metrics remain separate in the tables below.

The following tables cover general text translation, low-resource translation, and low-resource instruction following. FLORES uses COMET-22; WMT26 uses a judge score. instTrans reports translation quality and instruction following separately. MEME measures translation quality for community and cultural expressions. Higher is better for every metric in the first table below; scales differ across columns.

Model FLORES ↑ WMT26 ↑ instTrans quality ↑ instTrans IFscore ↑ MEME ↑
Index-Translate-35B-A3B (preview) 0.8794 76.76 0.6901 0.8336 0.7405
Index-Translate-9B 0.8789 75.35 0.6771 0.8209 0.7387
Index-Translate-2B 0.8655 60.26 0.5391 0.7569 0.6443
Hy-MT2-7B 0.8747 60.51 0.5143 0.6079 0.5139
Hy-MT2-30B-A3B 0.8787 66.81 0.5725 0.6415 0.5812
DeepSeek-V4.1-Flash 0.8762 83.55 0.6068 0.6374 0.7424
GPT-5.6-Sol 0.8650 89.10 0.6902 0.7624 0.7194
Gemini 3.5 Flash Lite 0.8750 79.52 0.6068 0.6374 0.7034

Low-resource translation and instruction following

FLORES_minor_pair evaluates general translation in low-resource languages; instTrans_minor reports translation quality and instruction following separately. Off-target is the percentage of outputs in a language other than the target language; lower is better. Higher is better for the other metrics. Bold marks the best value in each column of this table.

Model FLORES_minor_pair
COMET-22 ↑
FLORES_minor_pair
XCOMET-XXL ↑
FLORES_minor_pair
off-target ↓
instTrans_minor
Quality ↑
instTrans_minor
IFscore ↑
instTrans_minor
off-target ↓
Index-Translate-35B-A3B (preview) 0.8168 0.7164 2.4% 0.5151 0.7715 4.05%
Index-Translate-9B 0.7992 0.6805 4.0% 0.5222 0.7725 3.47%
Index-Translate-2B 0.7377 0.4817 4.2% 0.3050 0.6586 3.97%
Hy-MT2-7B 0.4626 0.3334 35.7% 0.1121 0.2405 45.40%
Hy-MT2-30B-A3B 0.6746 0.5359 14.5% 0.2246 0.4449 15.47%
DeepSeek-V4.1-Flash 0.8333 0.7297 1.3% 0.4793 0.5854 5.73%
GPT-5.6-Sol 0.7669 0.6918 10.4% 0.5757 0.6866 7.73%
Gemini 3.5 Flash Lite 0.8122 0.6995 3.3% 0.3927 0.5584 12.18%

Among the three Index-Translate models, 35B-A3B (preview) has the highest FLORES_minor_pair COMET-22 and XCOMET-XXL scores (0.8168 / 0.7164), with a 2.4% off-target rate. On instTrans_minor, Index-Translate-9B achieves the highest IFscore (0.7725) and lowest off-target rate (3.47%) among all compared models; its quality score is 0.5222.

Full tables retain all comparison models, low-resource metrics, WMT24++, IFMTBench, domain averages, general capabilities, speech, SandGlass, and long-document results. Detailed settings and analysis are in the technical report.

Index-Homura and Index-NativeLong

Index-Homura evaluation

SandGlass overall score and length adherence from the demo; the overall score differs from the separate translation-quality metric in the detailed tables.

Index-NativeLong 64K evaluation

GuoFeng and BWB Track A3 at 64K Chinese-side tokens.

For the specialized models, Index-Homura-9B reaches 81.92% within 10% of the target syllable count on SandGlass. Index-NativeLong-9B scores 0.7891 / 0.7683 / 0.8848 on GuoFeng / BWB / Books. The full tables include translation-quality tradeoffs and evaluation notes.

Benchmarks

Find the datasets and metadata in the 🤗 Index-Translate Benchmarks collection. Evaluation scripts, a pinned data downloader and run instructions are available in benchmarks/.

Benchmark Public release Dataset Evaluation code and guide
instTrans 3,000 translation instruction tasks; 10 constraint types 🤗 InstTrans-Bench insttrans
MEME 3,638 Chinese-to-English cases; 703 terms and 857 senses 🤗 Meme-Translation-Bench meme
SandGlass 3,600 cases: 300 subtitles × 4 target languages × 3 length budgets 🤗 Sandglass-Bench sandglass
NAtIveLong Metadata for 84 BWB and 84 GuoFeng long-document cases, evaluation code and acquisition instructions 🤗 NAtIveLong nativelong

NAtIveLong includes no source novels or reference translations: obtain official BWB data separately; GuoFeng provides a local reconstruction tool. The instTrans release does not include the 2,793 low-resource extension tasks in the evaluation tables. Original prompts are preserved; each directory's RELEASE_REVIEW.md records review changes and validation limits. Historical scores have not been rerun. Each benchmark has its own licensing terms; the repository's root license does not replace them.

See the full evaluation settings and IFMTBench preprocessing for other results.

Applications

  • Browser extension: translate web pages through a locally deployed model using Chrome, Edge, or Firefox.
  • Video dubbing pipeline: extract audio, separate vocals, segment speech, translate and dub, then align the result to the original video.

Community

Join the QQ group (960123527) for discussion and feedback. You can also scan the QR code below to join.

QR code to join QQ group 960123527

Papers and citation

Browse all three papers in the 🤗 Hugging Face Papers collection.

If you use this model family or its related methods, please cite the relevant papers:

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September},
  eprint={2609.40181},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2609.40181}
}

@misc{homura2026,
  author={Ziang Cui and Mengran Yu and Chenyu Shi and Yingxuan Shi and Tianjiao Li},
  title={HOMURA: Taming the Sand-Glass for Time-Constrained LLM Translation via Reinforcement Learning},
  year={2026},
  eprint={2601.10187},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2601.10187}
}

@misc{rival2025,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Yanjun Zhao and
          Xiaojing Liu and Qiang Zhang and Qi Zhang and Xuanjing Huang and Jiayin Wang},
  title={RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation},
  year={2025},
  eprint={2506.05070},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2506.05070}
}

License and feedback

Apache-2.0. Questions and feedback are welcome through GitHub Issues.

About

A Multilingual Translation Model Family

Topics

Resources

Stars

588 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages