arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24825cs.AI

面向大规模评估中项目附带内容自动相似度分析的双维度大语言模型框架

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

  • University of Arkansas(阿肯色大学)
  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

Jing Huang, Jihong Zhang, Hua-Hua Chang

AI总结:

本研究提出双维度LLM驱动的AISA框架,其衍生指标在心理测量学验证中表现更优,应用于CAT时可提升估计稳定性、减少偏差,优于传统度量,支持多种评估场景。

AI中文摘要:

大规模评估的快速扩展以及自动项目生成的日益普及,加剧了附带内容冗余的问题,即项目间无意重复出现诸如措辞或语境框架等与构念无关的元素。传统相似度度量指标如BLEU或余弦相似度,往往无法同时捕捉导致感知冗余的细微结构和语义层次。本研究提出一种由大语言模型(LLM)驱动的自动项目相似度分析(AISA)双维度框架,通过结构化分解和语义相关性来量化相似度。心理测量学验证表明,LLM衍生的指标与与构念无关的局部依赖指标更为一致,且比传统基于文本的度量产生更连贯的项目参数分组。该框架还通过在计算机化自适应测试(CAT)中的应用进行评估,模拟结果显示,在项目选择中纳入基于LLM的相似度约束,可在效率损失极小的情况下提高估计稳定性并减少偏差,表现优于基于传统度量的约束。这些发现凸显了LLM驱动的AISA在各类评估场景中支持可扩展题库管理、内容感知测试组卷以及体验敏感型自适应测试的潜力。

英文摘要:

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity, often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously. This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and yield more coherent item parameter groupings than traditional text-based measures. The framework is further evaluated through its application in Computerized Adaptive Testing (CAT). Simulations reveal that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias with minimal efficiency trade-offs, outperforming constraints based on conventional metrics. These findings highlight the potential of LLM-powered AISA to support scalable bank curation, content-aware test assembly, and experience-sensitive adaptive testing across diverse assessment contexts.

补充信息

↑