发表机构
University of Pisa; CNR - IIT; CNR - ISTI(比萨大学; 意大利国家研究委员会 - 信息与电信研究所; 意大利国家研究委员会 - 信息科学与技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
综述20项关于AI描述STEM视觉的研究,关注可及性和人机交互。分析目标视觉类型等多方面内容,指出从静态转向交互式多模态系统,同时存在事实不准确等挑战,最后概述人机交互关键研究方向。
AI 中文摘要
科学、技术、工程和数学(STEM)领域视觉数据的激增给盲人或视力障碍者带来了获取障碍。虽然人工智能(AI)的最新进展为生成STEM图像的文本描述提供了新机会,但研究格局分散,对实际用户的影响有限。本系统综述考察了20项关于基于AI描述STEM视觉的同行评审研究,特别关注可及性和人机交互。遵循PRISMA方法和基于ROBIS的偏倚风险评估,分析了目标STEM视觉类型、采用的AI和机器学习架构、采用的数据集和评估指标以及描述传递的交互方式。分析表明从静态、一次性替代文本转向集成对话界面、键盘导航和音频或触觉反馈的交互式多模态系统。然而,关键挑战依然存在,包括事实不准确和幻觉、与盲人和低视力用户共同设计的以可及性为先的数据集稀缺以及严重依赖自动文本重叠指标。综述最后概述了人机交互的关键研究方向,强调用户控制的详细程度、可解释和可验证的AI管道以及将可及描述工具集成到主流STEM创作和学习环境中。
英文摘要
The proliferation of visual data in Science, Technology, Engineering, and Mathematics (STEM) fields presents accessibility barrier for individuals with blindness or visual impairments. While recent advances in Artificial Intelligence (AI) offer new opportunities to generate textual descriptions of STEM images, the research landscape is fragmented and its impact on real users remains limited. This systematic survey examines 20 peer-reviewed studies on AI-based techniques for describing STEM visuals, with a specific focus on accessibility and human-computer interaction. Following the PRISMA methodology and a ROBIS-based risk-of-bias assessment, the review analyzes (i) the types of STEM visuals targeted, (ii) the AI and machine learning architectures employed, (iii) the datasets and evaluation metrics adopted, and (iv) the interaction modalities through which descriptions are delivered. The analysis reveals a shift from static, one-shot alt text toward interactive and multimodal systems that integrate conversational interfaces, keyboard navigation, and audio or haptic feedback. However, critical challenges persist, including factual inaccuracies and hallucinations, the scarcity of accessibility-first datasets co-designed with blind and low-vision users, and a heavy reliance on automatic text-overlap metrics that poorly capture perceived usefulness and trust. The survey concludes by outlining key research directions for HCI, emphasizing user-controlled verbosity, explainable and verifiable AI pipelines, and the integration of accessible description tools into mainstream STEM authoring and learning environments.
CommentsThis is the Author's Original Manuscript of an article accepted for publication in International Journal of Human-Computer Interaction, published by Taylor and Francis. The Version of Record is available at DOI 10.1080/10447318.2026.2668031
DOI:10.1080/10447318.2026.2668031