文章背景与核心概要
本文探讨了驱动大语言模型(LLM)中产生类人概念表征的计算因素。通过利用认知科学中的三元组相似度任务(基于THINGS数据库的概念),作者对75个以上的模型进行了评估,发现在众多因素中,指令微调(instruction fine-tuning)和更大的注意力头维度(attention head dimensionality)能够显著提升模型与人类表征的对齐程度。相比之下,参数规模、激活函数和多模态预训练对对齐的影响微乎其微。
此外,该研究揭示了当前的大语言模型基准在可靠测量人类概念对齐方面存在严重不足。这项研究不仅找出了将LLM推进为人类概念表征模型的核心计算要素,还填补了当前LLM评估方法中的关键空白,为未来兼具认知合理性与高性能的AI模型开发提供了重要指导。
揭示大语言模型中类人表征的计算要素
arXiv ID: 2510.01030 [cs.AI]
作者: Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh
提交历史:
* [v1] 2025年10月1日,周三
* [v2] 2026年8月31日(当前版本)
许可证: Creative Commons Attribution 4.0 (
查看许可证)
执行摘要
本文研究了驱动大语言模型(LLMs)中产生类人概念表征的计算因素。通过使用来自THINGS数据库的概念,并结合认知科学中的三元组相似度任务对75个以上的模型进行评估,作者发现指令微调和更大的注意力头维度能显著改善模型与人类的对齐程度。相反,参数规模、激活函数和多模态预训练的影响则微乎其微。此外,研究还表明,当前的许多LLM基准在可靠测量人类概念对齐方面大体上是不够充分的。
Executive Summary
This paper investigates the computational factors that drive human-like conceptual representations in Large Language Models (LLMs). By evaluating over 75 models using a cognitive science triplet similarity task with concepts from the THINGS database, the authors discover that instruction fine-tuning and larger attention head dimensionality significantly improve human-model alignment. Conversely, parameter size, activation functions, and multimodal pretraining have minimal impact. Furthermore, the study reveals that current LLM benchmarks are largely inadequate for reliably measuring conceptual alignment with humans.
摘要
人类将各种知觉和语言输入转化为结构化行为的能力,被认为依赖于学习强健的概念表征。基于Transformer的大语言模型(LLMs)的快速发展,涌现出了各种与模型构建相关的计算要素——包括架构、微调方法和训练数据集等,然而目前尚不清楚哪些要素对于开发类人概念表征最为关键。
此外,当前的大多数基准测试并不适合测量表征对齐,这使得LLM在这些测试中的得分无法可靠地评估它们是否在作为认知模型不断进步。为了解决这些局限性,我们评估了75个以上的模型在三元组相似度任务上的表现。该方法是认知科学中用于测量概念表征的成熟方法,采用的是THINGS数据库中的概念。
我们发现,指令微调和更大的注意力头维度是预测人类对齐的最强指标之一,而激活函数的选择、多模态预训练和参数规模对对齐的影响有限。对齐分数与现有基准分数之间的相关性表明,虽然某些基准(例如 BigBenchHard)比其他基准(例如 MUSR)能更好地捕捉表征对齐,但没有一个基准能完全解释人机对齐中的方差,这证明了它们的不足。综上所述,我们的研究结果突显了将LLM推进为人类概念表征模型的关键计算要素,并填补了LLM评估中的一个关键空白。
Abstract
The human ability to translate diverse perceptual and linguistic inputs into structured behavior has been thought to rest on learning robust representations of concepts. The rapid advancement of transformer-based large language models (LLMs) has surfaced a diversity of computational ingredients relevant for model building—architectures, fine-tuning methods, and training datasets among others—yet it remains unclear which are most crucial for developing human-like conceptual representations.
Further, most current benchmarks are ill-suited to measuring representational alignment, making LLMs' scores on them unreliable for assessing whether they are progressing as cognitive models. We address these limitations by evaluating over 75 models on a triplet similarity task, a method well established in cognitive science for measuring conceptual representations, using concepts from the THINGS database.
We find that instruction fine-tuning and larger attention head dimensionality are among the strongest predictors of human alignment, while activation function choice, multimodal pretraining, and parameter size have limited influence on alignment. Correlations between alignment scores and existing benchmark scores reveal that while some benchmarks (e.g., BigBenchHard) better capture representational alignment than others (e.g., MUSR), none fully accounts for the variance in human-model alignment, demonstrating their insufficiency. Taken together, our findings highlight key computational ingredients for advancing LLMs as models of human conceptual representation and address a key gap in LLM evaluation.
获取论文与资源
- 全文选项:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 外部书目与引用工具:
- NASA ADS
- Google Scholar
- Semantic Scholar
Access Paper & Resources
- Full-Text Options:
- View PDF
- HTML Version (Experimental)
- TeX Source
- External Bibliographic & Citation Tools:
- NASA ADS
- Google Scholar
- Semantic Scholar