发表机构
Australian National University(澳大利亚国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在基于视觉的简答题评分任务中,用边界框裁剪学生答案对小语言模型性能的影响,通过实验表明此方法能显著提高评分准确性并降低计算成本,是大规模视觉教育评估中部署小语言模型的关键预处理步骤。
AI 中文摘要
在教育环境中部署小语言模型(SLMs)在隐私、成本和可扩展性方面具有显著优势。然而,由于处理大图像的高计算成本和整页存在的视觉干扰,SLMs在诸如批改手写学生考试等复杂的基于视觉的任务中常常遇到困难。在本文中,我们研究了使用边界框裁剪学生答案是否可以提高SLMs在简答题评分任务中的准确性和计算效率。使用2025年澳大利亚物理奥林匹克竞赛的扫描手写答案数据集,我们在不同的思维链(CoT)提示和图像裁剪条件下评估了几个参数从4B到72B的模型的性能。我们的结果表明,使用边界框显著提高了评分准确性并降低了各模型的计算成本(FLOPs)。我们得出结论,边界框是在大规模基于视觉的教育评估中部署SLMs的关键预处理步骤。
英文摘要
The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms of privacy, cost, and scalability. However, SLMs often struggle with complex vision-based tasks, such as grading handwritten student exams, due to the high computational cost of processing large images and the visual distractions present on a full page. In this paper, we investigate whether cropping student responses using bounding boxes can improve the accuracy and computational efficiency of SLMs on a short-answer grading task. Using a dataset of scanned handwritten responses from the 2025 Australian Physics Olympiad, we evaluate the performance of several models ranging from 4B to 72B parameters under varying conditions of Chain of Thought (CoT) prompting and image cropping. Our results demonstrate that using bounding boxes significantly improves grading accuracy and reduces computational cost (FLOPs) across models. We conclude that bounding boxes are a crucial pre-processing step for deploying SLMs in large-scale, vision-based educational assessments.
CommentsAccepted for 1st Workshop on Small Language Models for Education (SLM4ED '26) at AIED 2026