跳转至

大规模测评中自动化题目偶然内容相似度分析的双维度大模型框架

文章背景与核心概要

随着大规模测评的飞速扩张和自动生成题目(Automatic Item Generation)技术的广泛应用,与构念无关的元素(如用词或语境框架)在题目之间无意中重复的问题,已成为一项严峻的挑战。传统的文本相似度指标(如 BLEU 或余弦相似度)往往难以捕捉细微的结构和语义冗余。

本文介绍了一种用于自动化题目相似度分析(AISA)的双维度大语言模型(LLM)框架,该框架通过结构化分解(Structured Decomposition)和语义相关性(Semantic Relatedness)来实现具体操作。心理计量学评估表明,与传统文本测量方法相比,LLM 衍生出的指标与构念无关局部依赖性(construct-irrelevant local dependence)指标具有更高的一致性,并能提供更优的题目参数分组。此外,计算机化自适应测试(CAT)中的模拟实验表明,集成基于 LLM 的相似度约束能够有效提高估计稳定性并减少偏差。


📄 执行摘要

With the rapid expansion of large-scale assessments and automatic item generation, unintended repetition of construct-irrelevant elements (such as wording or contextual framing) has become a critical challenge. Traditional text similarity metrics (like BLEU or cosine similarity) frequently fail to capture nuanced structural and semantic redundancy.

This paper introduces a Dual-Dimensional Large Language Model (LLM) Framework for Automated Item Similarity Analysis (AISA), operationalized through Structured Decomposition and Semantic Relatedness. Psychometric evaluations demonstrate that LLM-derived metrics align significantly better with indicators of construct-irrelevant local dependence and provide superior item parameter groupings compared to legacy text measures. Furthermore, simulations within Computerized Adaptive Testing (CAT) show that integrating LLM-based similarity constraints improves estimation stability and reduces bias efficiently.

随着大规模测评的快速扩张和自动生成题目技术的普及,人们对偶然内容冗余(incidental content redundancy)的担忧日益加剧。在这种现象中,与构念无关的元素(例如用词或语境框架)在不同题目之间无意中出现重复。诸如 BLEU 或余弦相似度等传统相似度指标,往往无法同时捕捉驱动感知冗余的微妙结构层和语义层。本研究提出了一种由大语言模型(LLM)驱动的自动化题目相似度分析(AISA)双维度框架,通过结构化分解和语义相关性来实现相似度运算。心理计量学验证表明,与传统的基于文本的测量方法相比,LLM 衍生出的指标与构念无关局部依赖性的指标吻合度更高,并能产生更具连贯性的题目参数分组。该框架通过在计算机化自适应测试(CAT)中的应用进行了进一步评估。模拟结果表明,将基于 LLM 的相似度约束纳入题目选择中,能够在将效率权衡降至最低的同时,提高估计稳定并减少偏差,其表现优于基于传统指标的约束。这些发现突显了由 LLM 驱动的 AISA 在支持可扩展题库管理、内容感知组题以及各种测评场景下体验敏感的自适应测试方面的潜力。


📋 元数据

  • arXiv ID: arXiv:2608.24825 [cs.AI]
  • 作者: Jing Huang, Jihong Zhang, Hua-Hua Chang
  • 提交时间: 2026年8月25日
  • 主要学科: 人工智能 (cs.AI)
  • 篇幅范围: 26页,6幅图表
  • DOI: 10.48550/arXiv.2608.24825

🔍 摘要

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity, often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously. This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and yield more coherent item parameter groupings than traditional text-based measures. The framework is further evaluated through its application in Computerized Adaptive Testing (CAT). Simulations reveal that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias with minimal efficiency trade-offs, outperforming constraints based on conventional metrics. These findings highlight the potential of LLM-powered AISA to support scalable bank curation, content-aware test assembly, and experience-sensitive adaptive testing across diverse assessment contexts.

随着大规模测评的快速扩张和自动生成题目技术的普及,人们对偶然内容冗余(incidental content redundancy)的担忧日益加剧。在这种现象中,与构念无关的元素(例如用词或语境框架)在不同题目之间无意中出现重复。诸如 BLEU 或余弦相似度等传统相似度指标,往往无法同时捕捉驱动感知冗余的微妙结构层和语义层。本研究提出了一种由大语言模型(LLM)驱动的自动化题目相似度分析(AISA)双维度框架,通过结构化分解和语义相关性来实现相似度运算。心理计量学验证表明,与传统的基于文本的测量方法相比,LLM 衍生出的指标与构念无关局部依赖性的指标吻合度更高,并能产生更具连贯性的题目参数分组。该框架通过在计算机化自适应测试(CAT)中的应用进行了进一步评估。模拟结果表明,将基于 LLM 的相似度约束纳入题目选择中,能够在将效率权衡降至最低的同时,提高估计稳定并减少偏差,其表现优于基于传统指标的约束。这些发现突显了由 LLM 驱动的 AISA 在支持可扩展题库管理、内容感知组题以及各种测评场景下体验敏感的自适应测试方面的潜力。


🔗 访问与全文资源