arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为题目难度预测表征视觉证据:视觉文本化与原生图像建模

Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling

Han Chen, Ming Li, Hong Jiao, Tianyi Zhou

arXiv 2608.04554首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence; University of Maryland(穆罕默德·本·扎耶德人工智能大学; 马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对题目难度预测的视觉证据表征问题,对比仅用文本、视觉文本化与原生图像建模三种方式,发现原生图像建模是有竞争力的替代方案,其有效性取决于VLMs的适配方式。

AI 中文摘要

在足够学生响应可用前,基于题目内容预测其难度可为新开发的题目提供初始评估。现有方法通常将题干和选项表征为文本。当数学题目包含视觉组件时,常见流程是先将该证据文本化,再应用文本预测器。本文研究问题:应如何为题目难度预测表征视觉证据?本文对比了仅用题目文本、将视觉证据转化为语言的视觉文本化、保留原始图像的原生图像建模这三种方式。使用经学生响应校准难度的Eedi题目,本文直接训练大型语言模型(LLMs)和视觉-语言模型(VLMs)进行难度回归。两种视觉界面均达到最低点估计,尽管领先系统无法可靠排序。Open-VLM文本化对所有评估的LLMs产生更低的RMSE点估计,而更广泛的适配对所有原生图像VLMs产生更低的RMSE点估计。测试时干预显示其依赖配对的完整题目图像,但无法分离额外视觉组件。两种视觉界面还会产生部分互补的题目级错误,且计算流程差异显著。因此,不应将文本化视为唯一实用界面:原生图像建模是一种有竞争力的替代方案,其有效性取决于VLMs的适配方式。

英文摘要

Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typically represent the question stem and answer choices as text. When mathematics items contain visual components, a common pipeline first textualizes that evidence and then applies a text predictor. We ask: how should visual evidence be represented for item difficulty prediction? We compare question text alone, visual textualization, which expresses visual evidence in language, and image-native modeling, which retains the original image. Using Eedi items with difficulty calibrated from student responses, we train large language models (LLMs) and vision-language models (VLMs) directly for difficulty regression. Both visual interfaces achieve the lowest point estimates, although the leading systems cannot be reliably ordered. Open-VLM textualization yields lower RMSE point estimates for all evaluated LLMs, while broader adaptation does so for all image-native VLMs. Test-time interventions show dependence on the paired full-item image, but do not isolate the additional visual component. The two visual interfaces also make partially complementary item-level errors and differ substantially in computational workflow. Thus, textualization should not be treated as the only practical interface: image-native modeling is a competitive alternative whose effectiveness depends on how the VLM is adapted.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑