arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpecialEduBench:面向自闭症儿童语言干预中知识、技能与态度的视觉语言模型基准测试

SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children

Jihoi Na, Taeyeong Kim, Sungjune Kong, Jaemin Jung, Min Joung Park, Kyungtae Joo, Ahhyun Kim, Shim Jaechang, Sooyoung Joo, Dongjin Ka, SeJoong Kim, Jimin Kim, HyunJin Jung, Unggi Lee

arXiv 2609.26090首次发表:更新:

AI 中文总结

针对自闭症儿童语言干预,提出SpecialEduBench基准,从知识、技能、态度三维度评估视觉语言模型教学能力,含4537个知识及200技能、68态度条目,发现前沿模型在压力情境下仍存在显著不足。

AI 中文摘要

语言是大多数自闭症儿童早期干预的目标。由于目标和方法因儿童而异,这项工作落在教师身上,教师一次只面对一个孩子,并在每个场景展开时做出判断。人工智能如今正被引入这项工作,然而面向特殊教育的基准测试所问的是模型知道什么,而非它在孩子面前做了什么。构建这样的基准并非易事,因为一个回应是否属于良好的教学取决于孩子刚刚做了什么,因此没有现成的标准答案可循。决定其优劣的证据既有视觉性的也有言语性的,因为等待的时长、目光的转移以及孩子的领会都不会在文字记录中留下痕迹。我们引入了\emph{SpecialEduBench},它从知识、技能和态度三个维度衡量教学能力,包含4,537个知识条目,以及基于录制的干预过程构建的200个技能条目和68个态度条目,其中态度条目将压力与监控交叉组合成192个响应单元。七位特殊教育专家编写、评分并审查了这些条目,我们根据他们设定的参考分数修订了评判模型的指令。在八个前沿视觉语言模型上,没有任何一个维度达到饱和,因为最强的模型在诚实性单元上仍有约十分之一的失败。模型在事实性知识上表现趋同,而在情境性任务上表现分化,失败集中在施加压力的场景中。我们期望该基准能在部署前作为审计工具,并作为针对该领域构建模型的起点。

英文摘要

Language is the target of most early intervention for autistic children. Because the goal and the method change from child to child, the work falls to a teacher who takes one child at a time and judges each scene as it unfolds. Artificial intelligence is now being brought to that work, yet the benchmarks that reach special education ask what a model knows rather than what it does in front of a child. Building one is not straightforward, since whether a response is good teaching depends on what the child has just done, so no answer key applies. The evidence that settles it is visual as much as verbal, since the length of a wait, a shift of gaze, and the child's uptake leave no trace in a transcript. We introduce \emph{SpecialEduBench}, which measures pedagogical competence along knowledge, skill, and attitude, with 4,537 knowledge items and with 200 skill items and 68 attitude items built on recorded intervention, the attitude items crossing pressure with monitoring into 192 response cells. Seven special-education experts wrote, scored, and reviewed the items, and we revised the judge model's instruction against the reference scores they set. Across eight frontier vision-language models no axis is saturated, since the strongest still fails about a tenth of the honesty cells. The models converge where the knowledge is factual and separate where the task is situated, and the failures gather where pressure is applied. We intend the benchmark as an audit to run before deployment and as a starting point for models built for this domain.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑