评估面向无障碍灾害援助的跨文本与音频模态多模态大语言模型
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance
中文总结 AI 辅助
本文评估了开放权重多模态大语言模型在灾害援助场景下跨文本与音频模态的响应一致性,发现其存在模态依赖的不公平性,为构建公平的灾害沟通AI工具提供了设计建议。
中文摘要 AI 辅助
有效的灾害风险沟通是一项基础性人道主义挑战,但当前应急基础设施无法满足有使用和功能需求的个体的需求,包括听障人士、孕妇、带幼儿的母亲以及患有痴呆症的老年人。人工智能(AI)的最新进展,尤其是多模态大语言模型(MM-LLMs),展现出在单个统一系统(如聊天机器人)中服务文本、音频、图像和视频等不同模态下多样化用户的强大能力。然而,它们的适用性取决于一个未得到充分审查的特性:无论用户通过何种模态进行沟通,这些系统是否能产生一致、可操作的输出。在本文中,我们使用涉及四个不同弱势角色的真实紧急警报场景,对开放权重MM-LLMs的现状进行了全面分析。当给定相同任务场景时,我们评估这些最先进(SOTA)模型在文本和音频模态下响应的一致性。研究结果表明,没有任何模型能在不同模态间实现可靠的一致性,且对于有使用需求的角色,性能差距会进一步扩大,这会引入依赖模态的不公平性,损害这些系统的人道主义价值。这些结果为构建公平、可信且包容的灾害风险沟通AI工具提供了具体的设计建议。
英文摘要
Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. Recent advancements in Artificial Intelligence (AI), especially Multi-Modal Large Language Models (MM-LLMs), demonstrate powerful capabilities to serve diverse users across text, audio, image, and video modalities within a single unified system, such as a chatbot. However, their suitability for deployment rests on a property that receives limited scrutiny, i.e., whether these systems produce consistent, actionable outputs regardless of the modality through which a user communicates. In this paper, we conduct a comprehensive analysis to understand the status of open-weight MM-LLMs using real emergency alert scenarios across four different vulnerable personas. These state-of-the-art (SOTA) models are evaluated on consistency of responses across text and audio modalities when the same task scenario is given. Findings indicate that no model achieves reliable consistency across modalities, and that performance gaps are heightened for personas with access needs, introducing modality-dependent inequity that undermines the humanitarian value of these systems. These results inform concrete design recommendations for building equitable, trustworthy, and inclusive AI tools for disaster risk communication.