跳转至

文章背景与核心概要

随着大语言模型(LLM)在历史语言翻译中展现出惊人的潜力,它们在数字人文(Digital Humanities)工作流中的应用日益广泛。然而,缺乏可靠的自动评估手段严重限制了其深入应用。本文以“文言文到英文”(Classical Chinese to English)的翻译为主要测试案例,深入探讨了最初为现代语言开发的现有自动评估指标是否能够可靠地检测历史语言中的翻译错误。

研究团队引入了一个基于“最小对”(minimal pairs)的诊断框架,以捕捉学术应用中突出的错误类型,并测试了基于参考译文(reference-based)和无参考译文(reference-free)的指标对错误敏感度以及对合规文本变体的容忍度。研究结果表明,所有现有的评估指标都表现出明显的盲区,其中 MetricX24 的综合表现最佳。这项研究凸显了为具有历史和文化独特性质的翻译场景开发更强大、更具可解释性指标的迫切需求。


Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

评估指标能检测出文言文到英文翻译中的错误吗?

Summary

总结

This research paper investigates whether existing automatic evaluation metrics, originally developed for modern languages, can reliably detect translation errors when applied to historical languages. Using the translation of Classical Chinese to English as a primary test case, the authors introduce a diagnostic framework based on minimal pairs to evaluate error sensitivity and tolerance to valid textual variation. The study reveals that while all current metrics exhibit distinct blind spots, MetricX24 performs best overall, highlighting a pressing need for more robust, interpretable metrics tailored for historically and culturally distinct translation settings.

本研究论文探讨了最初为现代语言开发的现有自动评估指标,在应用于历史语言时是否能可靠地检测出翻译错误。以文言文到英文的翻译作为主要测试案例,作者引入了一个基于最小对(minimal pairs)的诊断框架,用以评估指标对错误的敏感度以及对有效文本变体的容忍度。研究表明,虽然所有当前的指标都表现出明显的盲区,但 MetricX24 的整体表现最好,这凸显了为具有历史和文化独特性的翻译场景定制更强健、更具可解释性指标的迫切需求。


Document Metadata

文档元数据

  • arXiv ID: arXiv:2608.08283 [cs.CL]
  • Primary Subject: Computation and Language (cs.CL)
    • 主要学科: 计算与语言 (cs.CL)
  • Secondary Subjects: Artificial Intelligence (cs.AI)
    • 次要学科: 人工智能 (cs.AI)
  • Publication History:
  • Submitted on August 8, 2026 (v1)
  • Last revised August 12, 2026 (v2)
    • 出版历史:
  • 2026年8月8日提交 (v1)
  • 2026年8月12日最后修订 (v2)
  • DOI: 10.48550/arXiv.2608.08283

Authors

作者

  • Osvaldo Quinjica
  • Eric Bennett
  • Xinchen Yang
  • Andrew Schonebaum
  • Marine Carpuat
    • Osvaldo Quinjica
    • Eric Bennett
    • Xinchen Yang
    • Andrew Schonebaum
    • Marine Carpuat

Abstract

摘要

Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.

尽管大语言模型在翻译某些历史语言时表现得令人惊叹,但由于缺乏可靠的评估手段,它们在数字人文工作流中的实用性受到了限制。我们以文言文到英文的翻译为例,研究了为现代语言开发的现有自动评估指标在此场景下是否可靠。我们引入了一个基于最小对的诊断框架,以捕捉学术应用中突出的错误类型,并对基于参考和无参考的指标进行了错误敏感度及有效变体容忍度的探测。我们发现所有指标都存在盲区,不过 MetricX24 的整体表现最好。我们的研究结果凸显了针对历史和文化独特翻译场景开发更强健、更具可解释性指标的必要性。


全文与资源链接