评估工具对人类和大型语言模型(LLM)测量的是同一事物吗?一项潜在结构分析
文章背景与核心概要
随着大型语言模型(LLM)的快速发展,研究人员经常使用原本为人类设计的教育评估工具(如标准化考试)来衡量AI的能力。这种做法隐含了一个假设:即模型在这些测试中的优异表现,反映了与人类相同的潜在能力。
本研究从心理测量学效度的视角出发,对这一假设提出了质疑。通过对比人类与六种多模态LLM在高中化学和大学入学考试定量推理部分的表现,研究人员利用探索性因子分析、因子一致性检验和重采样技术,分析了这些评估工具的潜在结构。研究发现,人类与LLM在因子结构上存在系统性差异,这表明现有的教育评估工具可能并未测量到AI与人类相同的能力维度,从而挑战了将人类教育评估直接用于衡量AI能力的有效性。
文章详情
摘要
大型语言模型的快速发展和广泛部署,使得理解其能力变得日益重要。一种常见的方法是使用最初为衡量人类技能和能力而设计的评估工具(如标准化考试)来评估LLM,并将这些工具上的表现作为证据,对LLM在评估所针对的相同技能上的潜在能力做出普遍性结论。
然而,从效度角度来看,此类推断要求在人类身上建立的“观察表现”与“潜在构念”之间的关系同样适用于LLM。特别是,转移分数解释的一个必要条件是评估响应的潜在结构必须具有相似性。在本研究中,我们考察了这一条件在两个教育背景下是否成立:高中化学和大学入学考试的定量推理部分。通过案例研究设计,我们比较了人类的响应数据与六种多模态LLM生成的响应。我们的分析方法结合了探索性因子分析、因子一致性检验和重采样,以评估人类学习者与LLM之间潜在结构的相似性。
在两项评估工具中,我们均发现人类与LLM的因子结构之间存在系统性差异,这表明所分析的评估工具对于人类和LLM而言,可能并未测量相同的构念。这些发现对那些使用教育评估来断言AI能力的评估实践的有效性提出了质疑。
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the assessments are intended to measure in humans.
However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs.
Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.
元数据
- 学科分类: 人机交互 (
cs.HC);人工智能 (cs.AI);计算语言学 (cs.CL) - DOI: 10.48550/arXiv.2608.15630
- 许可协议: 知识共享 署名-非商业性使用-禁止演绎 4.0 国际 (许可图标:
)
- Subjects: Human-Computer Interaction (
cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)- DOI: 10.48550/arXiv.2608.15630
- License: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (License icon:
)
访问链接与资源
- 全文选项:
- 查看 PDF
- HTML 版本 (实验性)
- TeX 源码
- 外部引用与工具:
- Google Scholar
- Semantic Scholar
- NASA ADS
- Full-Text Options:
- View PDF
- HTML Version (Experimental)
- TeX Source
- External Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS