arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用人工智能查询多模态科学论文:盲人、低视力和视力正常科学家的实践与偏好

Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists

Arnavi Chheda-Kothary, Lucy Lu Wang, Joseph Chee Chang, Jonathan Bragg

arXiv 2607.18514首次发表:更新:

发表机构

Paul G. Allen School of Computer Science \& Engineering University of Washington Seattle WA USA; The Information School University of Washington Allen Institute for AI Seattle WA USA; Allen Institute for AI Seattle WA USA; University of Washington; Allen Institute for AI(; ; ; ; )

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究不同视力科学家如何用ChatGPT和Gemini查询多模态科学文档,通过采访了解其审查多模态内容的实践、对AI响应的反馈等,发现相关问题,贡献数据集,讨论对AI科学QA系统的影响。

AI 中文摘要

视觉图表、图形和表格是科学论文的核心,传达着文本之外的信息。传统上,盲人或低视力(BLV)科学家依靠静态替代文本获取论文中的图形,而人工智能(AI)的兴起使交互式问答(QA)成为视觉探索的可行范式,但科学家如何在实践中使用视觉QA以及如何提高其可访问性却鲜为人知。在这项工作中,我们采访了来自不同STEM领域的五位BLV科学家和五位视力正常的科学家,以了解他们如何使用ChatGPT和Gemini这两种人工智能工具查询多模态科学文档。我们的研究结果描述了科学家如何审查多模态内容,包括与视觉内容互动的现有实践(以及可访问性变通方法),以及对人工智能生成的多模态查询响应适用性的反馈。我们还发现,模糊或不完整的图像描述以及更广泛的错误人工智能输出,可能导致BLV和视力正常的科学家放弃人工智能工作流程。为支持未来研究,我们还贡献了一个包含115个查询和响应的数据集,这些查询和响应来自参与者在其领域与人工智能工具对论文的交互。最后,我们讨论了对人工智能驱动的科学QA系统的影响,并强调了跨能力和领域访问的考虑因素。

英文摘要

Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind or low-vision (BLV) scientists have traditionally relied on static alternative text to access figures in papers, the rise of artificial intelligence (AI) has made interactive question-answering (QA) a feasible paradigm for visual exploration; yet little is known about how scientists use visual QA in practice or how to improve its accessibility. In this work, we interview five BLV and five sighted scientists across different STEM fields to understand how they use two AI tools, ChatGPT and Gemini, to query multimodal scientific documents. Our findings characterize how scientists review multimodal content, including existing practices (along with accessibility workarounds) for engaging with visuals, and feedback on the suitability of AI-generated responses to multimodal queries. We further find that vague or incomplete image descriptions, as well as incorrect AI outputs more broadly, can cause both BLV and sighted scientists to abandon AI workflows. To support future research, we additionally contribute a dataset of 115 queries and responses from our participants' interactions with the AI tools for papers in their field. We close by discussing implications for AI-powered scientific QA systems, emphasizing considerations for access across abilities and domains.

DOI:10.1145/3797867.3829024

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑