论文概要
研究领域: NLP 作者: Vilém Zouhar, Niyati Bafna, Mukund Choudhary 发布时间: 2026-09-06 arXiv: 2509.04283
中文摘要
为了科学进步,我们需要能够测试最先进模型极限的基准,以及能告诉我们失败案例的评估方法。随着模型变得更强,机器翻译的标准基准正趋于饱和。此外,自动翻译指标不可靠、易受奖励黑客攻击,且提供无法操作的评估。即使黄金人类评估也并非没有问题,因为它常常缺乏可重复性、客观性和可扩展性。总体而言,这阻碍了我们跟踪该领域的客观进展和识别改进路径。我们推出了Last Translation Benchmark,这是一个由人类撰写并经过同行评审的示例集合(文本、图像、音频、视频),旨在打破领先的机器翻译模型。我们还提出了一种新的评估方法:每个示例都附带手工制作的验证规则,描述该示例上的具体失败案例,从而实现可靠且可操作的评估。Last Translation Benchmark是一个接受持续贡献的动态数据集。最新版本为LTBv1,包含2026年9月1日前接受的贡献,并计划随着新数据的持续收集而发布更新。
原文摘要
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present …
— 自动采集于 2026-09-07
#论文 #arXiv #NLP #小凯
