Official repository of "Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation"
We introduce MENT (Meta-Evaluation dataset of Non-Literal Translation), a human-annotated meta-evaluation dataset to systematically assess MT evaluation metrics.
Download: You can download the dataset from .
Our comprehensive meta-evaluation reveals the unreliability of MT metrics on non-literal content. Traditional metrics are fundamentally limited by the lack of deep semantic understanding, while LLM-as-a-Judge paradigms are hindered by the static knowledge cutoff and inherent score inconsistency.
The scripts of meta-evaluation are based on MT-Metrics-Eval (MTME). All scripts are located in meta_eval/.
We explore the agentic MT evaluation framework RATE (Reflective Agentic Translation Evaluation), architected around a centralized Core Agent, and orchestrates three functional sub-agents: the Evaluation Agent for pointwise assessment, the Search Agent for online knowledge retrieval, and the Comparison Agent for calibration by pairwise evaluation.
The core implementation of the framework is located at agentic/ directory.
@misc{tian2026literalmappingbenchmarkingimproving,
title={Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation},
author={Yanzhi Tian and Cunxiang Wang and Zeming Liu and Heyan Huang and Wenbo Yu and Dawei Song and Jie Tang and Yuhang Guo},
year={2026},
eprint={2601.07338},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2601.07338},
}


