(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
Minghao Wu1, Yulin Yuan2, Gholamreza Haffari1, Longyue Wang3 1Monash University 2University of Macau 3Tencent AI Lab Longyue Wang is the corresponding author:
Abstract
Recent advancements in machine translation (MT) have significantly enhanced translation quality across various domains. However, the translation of literary texts remains a formidable challenge due to their complex language, figurative expressions, and cultural nuances. In this work, we introduce a novel multi-agent framework based on large language models (LLMs) for literary translation, implemented as a company called TransAgents, which mirrors traditional translation publication process by leveraging the collective capabilities of multiple agents, to address the intricate demands of translating literary works. To evaluate the effectiveness of our system, we propose two innovative evaluation strategies: Monolingual Human Preference (MHP) and Bilingual LLM Preference (BLP). MHP assesses translations from the perspective of monolingual readers of the target language, while BLP uses advanced LLMs to compare translations directly with the original texts. Empirical findings indicate that despite lower $d$ -BLEU scores, translations from TransAgents are preferred by both human evaluators and LLMs over human-written references, particularly in genres requiring domain-specific knowledge.
中文速览
文学翻译长期是机器翻译领域最难啃的硬骨头——复杂的修辞、文化典故和独特文风都让现有系统力不从心。为此,研究者搭建了一个叫 TransAgents 的多智能体(multi-agent)虚拟翻译公司,让 CEO、资深编辑、译者、本地化专家和校对等角色各司其职、协同完成整本书的翻译,模拟真实出版流程。由于 BLEU 等传统指标难以公正评价文学译作,他们还提出了两种新评估方式:让目标语言母语读者在不看原文的情况下选出更好的译文(单语人类偏好,MHP),以及让 GPT-4 对照原文直接比较两个译本(双语大模型偏好,BLP)。实验结果显示,TransAgents 的 d-BLEU 分数虽然垫底,却在人类读者和大模型评审眼中都优于人工参考译文,尤其在涉及历史背景和文化细节的题材上表现突出,同时翻译成本仅为专业人工翻译的约八十分之一——这说明用 BLEU 来衡量文学翻译质量本身就存在严重局限,而多智能体协作路线在文学翻译这一"最后前沿"上展现出了真实的应用潜力。
原文 arXiv:2405.11804;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2405.11804v2