为什么所有大语言模型都对日本文化“情有独钟”?解析大模型中隐藏的文化与地域偏见
文章背景与核心概要
大语言模型(LLM)在跨文化应用中的表现日益受到关注,但其内部往往潜藏着复杂的文化与地域偏见。本文探讨了当大模型面对通用、与文化相关的问题时,其所展现出的特定地域偏好。作者通过构建一个包含24种语言的多语言数据集——文化相关开放式问题(CROQ)分类体系,对主流大模型进行了深入评测。
研究发现了一个令人惊讶的全局性偏见:大模型在回答时表现出强烈的倾向,特别容易将答案导向特定的国家——其中最显著的就是日本。此外,输入语言的选择也会显著影响偏见的形式:英语等高资源语言通常会引发更多样化的输出,而低资源语言则倾向于突出该语言作为官方语言的国家。通过对模型训练动态的追踪,研究表明这些文化倾向主要产生于监督微调(SFT)阶段,而非初始的预训练阶段。

摘要 (Summary)
This research paper investigates the hidden cultural and regional biases embedded within Large Language Models (LLMs). While previous studies have evaluated LLM cultural capabilities, this work specifically examines regional preferences when answering generic, culture-related questions.
To conduct this evaluation, the authors introduce a new multilingual dataset based on a taxonomy of Culture-Related Open Questions (CROQ) across 24 languages. The findings reveal a surprising overarching bias: LLMs exhibit a pronounced tendency to gravitate toward specific countries—most notably Japan—in their responses.
Furthermore, the study highlights how input language impacts bias: * High-resource languages (like English) tend to elicit more diverse outputs. * Low-resource languages show strong inclinations toward highlighting countries where the input language holds official status. * Training dynamics: Investigation into the emergence of this bias suggests that these cultural slants primarily manifest during the supervised fine-tuning stage rather than initial pre-training.
本研究论文探讨了嵌入在大型语言模型(LLM)中的隐藏文化与地域偏见。尽管先前的研究评估过LLM的文化能力,但本工作专门检查了其在回答通用、与文化相关的问题时的地域偏好。
为了进行这一评估,作者引入了一个基于24种语言的“文化相关开放式问题(CROQ)”分类体系的新多语言数据集。研究结果揭示了一个令人惊讶的总体偏见:LLM在回答中表现出明显的倾向,即向特定国家(最显著的是日本)靠拢。
此外,该研究强调了输入语言如何影响偏见: * 高资源语言(如英语)往往会引发更多样化的输出。 * 低资源语言则表现出强烈的倾向,倾向于突出显示该输入语言为其官方语言的国家。 * 训练动态: 对这种偏见出现的机制的调查表明,这些文化倾向主要体现在监督微调阶段,而不是初始的预训练阶段。
文档元数据 (Document Metadata)
- arXiv ID: arXiv:2604.21751 [cs.CL]
- 主要学科: 计算与语言 (
cs.CL) - 次要学科: 人工智能 (
cs.AI);计算机与社会 (cs.CY) - 作者:
- Joseba Fernandez de Landa
- Carla Perez-Almendros
- Jose Camacho-Collados
- 提交历史:
- 提交于 2026年4月23日 (
v1) - 最后修订于 2026年8月28日 (
v2) - 数据集可用性: Hugging Face 数据集 (HiTZ/CROQ)
摘要原文 (Abstract)
LLMs have limitations when it comes to cultural coverage and competence, and in some cases, show specific cultural biases. Although prior studies have examined the cultural capabilities of LLMs, none have specifically investigated their regional preferences in generic culture-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ), with questions available in 24 languages. We evaluate LLMs by prompting them to answer questions from CROQ and provide a sample location. The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan in their answers. Moreover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs. Low-resource languages, on the other hand, show more inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training.
LLMs在文化覆盖范围和能力方面存在局限性,在某些情况下还会表现出特定的文化偏见。尽管先前的研究已经考察了LLM的文化能力,但没有一项研究专门调查其在通用文化相关问题中的地域偏好。在这项工作中,我们提出了一个基于文化相关开放式问题(CROQ)综合分类法的新数据集,其中包含24种语言的问题。我们通过提示LLM回答CROQ中的问题并提供样本位置来评估它们。结果表明,与以往的文化偏见研究相反,LLM在回答中表现出对日本等国家的明显偏好。此外,我们的结果表明,当使用英语或其他高资源语言进行提示时,LLM往往会提供更多样化的输出。另一方面,低资源语言表现出更多倾向于回答突出显示输入语言为其官方语言的国家的倾向。最后,我们还调查了这种文化偏见在LLM训练的哪个阶段出现,我们的结果表明,第一个明显的迹象出现在监督微调之后,而不是在预训练期间。
附加资源与访问链接 (Additional Resources & Access Links)
- 全文选项:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 外部引用与工具:
- Google 学术
- Semantic Scholar
- NASA ADS
- 社区书签:
- BibSonomy