arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00868cs.CVcs.HC

MIDAL:用于无障碍学习的数学图像描述

MIDAL: A Dataset of Math Image Descriptions for Accessible Learning

Rebeka Popek, Vaghawan Ojha, Young Hwan You

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出含2020张多级别数学图像的MIDAL数据集,用于训练视觉语言模型生成无障碍图像描述,还可微调语言模型提升数学推理能力,助力高等教育STEM内容无障碍性建设。

中文摘要 AI 辅助

许多开放教育资源缺乏无障碍性,尤其是深度图像描述。但在科学、数学等学科中,图像描述的撰写尤为困难,因为其包含大量复杂表达式和名称,且需匹配不同课程级别。为在一定程度上填补这一空白,我们推出了Math Image Descriptions for Accessible Learning(MIDAL),这是一个包含2020张数学图像的数据集,覆盖多个教育级别,用于训练视觉语言模型,使其能按照无障碍最佳实践生成图像描述。我们希望MIDAL能成为推动高等教育STEM内容无障碍性相关交流与创新的宝贵资源。该数据集不仅限于数学描述生成,还可用于微调语言模型,提升其数学推理与作答能力。

英文摘要

Many open educational resources are lacking in accessibility, especially in-depth image descriptions. In subjects like Science and Mathematics, however, it can be particularly difficult to write image descriptions since there can be many complicated expressions and names depending upon the course level. To help fill that gap in a small way, we introduce Math Image Descriptions for Accessible Learning (MIDAL), a math image-description dataset of 2,020 mathematical images spanning multiple educational levels, to aid in training vision language models to create image descriptions following accessibility best practices. We hope MIDAL is a valuable resource in enhancing the conversation and innovation regarding accessibility of STEM content in higher education. This dataset is however not just limited in math description generation but can also be used to fine-tune language models that can have improved mathematical reasoning and answers.

发表机构

  • E.K. Solutions Pvt. Ltd.(E.K.解决方案私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑