发表机构
Smarter Balanced, University of California-Santa Cruz; University of Minnesota-Twin Cities(加州大学圣克鲁兹分校 Smarter Balanced; 明尼苏达大学双城分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用大规模标准化测试的项目数据,通过融合DeBERTa分类器与Qwen3生成的项目评语,构建AIE模型预测项目接受/拒绝,可减轻人工评审负担,但对公平性相关项目的识别仍需人工评审。
AI 中文摘要
自动化项目评估(AIE)指的是使用计算方法评估项目质量,无需人工专家评审或对被评估项目进行实地测试。本研究旨在通过大规模标准化测试项目的历史拒绝数据,从项目文本中预测项目的接受与拒绝,以构建近乎全面的AIE模型。数据集包含52759个英语语言艺术(ELA)和数学项目,其中34%被永久拒绝用于未来的实际使用。拒绝原因包括心理测量特性不佳、内容问题、偏见与敏感性担忧以及非内容问题。我们在原始项目文本上微调了DeBERTaV3-large分类器,在Qwen3生成的项目评语上微调了第二个DeBERTa分类器,并构建了结合两者表示的融合模型。融合模型取得了最强的整体性能(准确率=0.75,F1值=0.64,AUC=0.80,灵敏度=0.64,特异性=0.81)。数学项目的预测(F1值=0.73,AUC=0.86)比ELA项目(F1值=0.51,AUC=0.72)准确得多。将决策阈值从0.5降至0.25,使ELA和数学项目的平均灵敏度分别提高至0.88和0.91,而特异性分别降至0.31和0.56,这在自动化项目生成场景中可能更可取,因为生成项目的成本低于评估项目。结合项目评语与原始项目文本在大多数拒绝原因上提升了性能。模型为难度更高的项目分配了更高的拒绝概率。然而,融合模型难以识别标记为偏见、敏感性、公平性或可访问性的项目,尤其是对于ELA项目。这些发现表明,基于文本的AIE在某些领域是可行的,可作为减轻人工评审和实地测试负担的实用工具,同时强调了对存在公平性担忧的项目进行人工评审的重要性。
英文摘要
Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item acceptance and rejection from item text using historical rejection data from a large-scale standardized testing program. The dataset contained 52,759 English language arts (ELA) and mathematics items with 34% permanently rejected from future operational use. Rejection reasons included poor psychometric properties, content issues, bias and sensitivity concerns, and non-content issues. We fine-tuned a DeBERTaV3-large classifier on raw item text, a second DeBERTa classifier on Qwen3-generated item critiques, and a fusion model combining representations from both. The fusion model achieved the strongest overall performance (Accuracy = .75, F1 = .64, AUC = .80, Sensitivity = .64, Specificity = .81). Prediction for math (F1 = .73, AUC = .86) was considerably more accurate than ELA (F1 = .51, AUC = .72). Lowering the decision threshold from .5 to .25 raised average sensitivity for ELA and math to .88 and .91, while reducing specificity to .31 and .56, respectively, which may be preferable in automated item generation contexts where generating items is cheaper than evaluating them. Incorporating item critiques alongside raw item text improved performance across most rejection reasons. The model assigned higher rejection probabilities to more difficult items. However, the fusion model struggled to identify items flagged for bias, sensitivity, fairness, or accessibility, especially for ELA. These findings suggest that text-based AIE is feasible in some areas and may offer a practical tool for reducing the burden of manual review and field testing, while also underscoring the importance of human review for items with fairness concerns.
Comments19 pages, 4 figures, 3 tables