跳转至

相似性贯穿始终:大语言模型的多语言泛化依赖于语言层面的相似性结构

文章背景与核心概要

随着大语言模型(LLM)在多任务处理能力上的不断提升,其在多语言环境、非英语语种以及低资源领域中的泛化能力仍是一个亟待解决的重大挑战。本文借鉴认知科学的研究视角,探讨了模型成功实现多语言泛化的内在机制,即这种泛化能力是否源于模型在相似性空间中构建了恰当的表征。

研究发现,大语言模型的潜在表征能够成功还原印欧语系的层级结构,将属于同一语支的语言在表征空间中紧密聚类。更重要的是,这种语言层面的相似性结构与模型在 XNLI(跨语言自然语言推理)基准测试上的表现呈现出显著的正相关性。这一结论揭示了模型泛化能力的本质:能够将相似语言进行相似表征的模型,在跨语言知识迁移方面具有天然的优势。


论文元数据

  • 作者: Supantho Rakshit, Adele Goldberg, Henry Conklin
  • 学科: 人工智能 (cs.AI); 计算与语言 (cs.CL)
  • arXiv ID: arXiv:2607.22699 [cs.AI]
  • DOI: 10.48550/arXiv.2607.22699
  • 会议: 已被 CogSci 2026(第48届认知科学学会年会,里约热内卢)接收为口头报告。
  • 提交日期: 2026年7月18日提交;2026年8月13日最后修订。

摘要

随着大语言模型(LLM)在各种任务中表现出越来越强的能力,其泛化能力(或缺乏泛化能力)在有限领域之外仍然难以量化且理解不足。特别是,众所周知,大语言模型在多语言泛化、非英语语言以及训练数据中覆盖不足的语言方面表现吃力。

为了理解其中的原因,并探究为何某些模型比其他模型表现更好,我们转向了认知科学领域长期的研究成果,论证了成功的泛化源于相似性空间中恰当的表征。我们研究了大语言模型的表征在多大程度上捕捉到了不同语言之间层级的相似性结构。

令人惊讶的是,我们证明了大语言模型的潜在表征在很大程度上还原了印欧语系的层级结构——将属于同一语支的语言在表征空间中紧密地组合在一起。此外,我们还证明了模型反映语言相似性结构的程度与其在 XNLI(一个多语言自然语言推理基准)上的表现相关。这扩展了关于大规模相似性驱动泛化的经典研究,表明那些将相似语言进行相似表征的模型,能够更好地实现从一种语言到另一种语言的知识泛化。

As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data.

To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs' representations capture the hierarchical similarity structure between distinct languages.

Strikingly, we show LLMs' latent representations largely recover the hierarchical structure of the Indo-European language family tree—grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.


访问链接与资源