arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27135cs.CL

语音与文本的不同解读:多模态模型中的跨模态不稳定性

Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models

Basel Mousi, Fahim Dalvi, Shammur Chowdhury, Firoj Alam, Nadir Durrani

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多模态模型在语音与文本、英语与阿拉伯语间的判断一致性问题,构建含10150张图像的基准,发现模态与语言转换会引入未被整体准确率捕捉的不一致性,语音还会放大部分失败,且公开了该基准。

中文摘要 AI 辅助

多模态基础模型越来越多地用于语音优先助手,这类助手需要解读口语查询并做出基于视觉的决策。然而,语义等价的查询是否会在模态(文本与语音)和语言(英语与阿拉伯语)之间产生一致的判断,目前仍不清楚。我们推出了一个语音增强的基于视觉的对比三元组基准,涵盖来自18个中东和北非(MENA)国家的10150张具有文化背景的图像,每张图像搭配一条受支持的陈述和两条看似合理但未受支持的替代陈述。我们将对比不稳定性定义为模型无法解决三元组中所有陈述的条件概率,该指标可将碎片化推理与完全失败区分开。我们在英语和阿拉伯语的文本与语音模态下评估了近期的多模态模型,发现模态和语言转换会引入大量三元组层面的不一致性,而这种不一致性并未被整体准确率完全捕捉,语音会放大部分失败情况。我们将该基准公开提供给研究界使用。

英文摘要

Multimodal foundation models are increasingly used in speech-first assistants that must interpret spoken queries and produce visually grounded decisions. Yet it remains unclear whether semantically equivalent queries yield consistent judgments across modality (text vs. speech) and language (English vs. Arabic). We introduce a speech-augmented visually grounded contrastive triplet benchmark spanning 10,150 culturally grounded images from 18 MENA countries, where each image is paired with one supported statement and two plausible but unsupported alternatives. We define contrastive instability as the conditional rate at which a model fails to resolve all statements within a triplet, isolating fragmented reasoning from complete failure. Evaluating recent multimodal models under text and speech in English and Arabic, we find that modality and language shifts introduce substantial triplet-level inconsistencies that are not fully captured by aggregate accuracy, with speech amplifying partial failures. We make the benchmark publicly available to the community.

发表机构

  • Qatar Computing Research Institute, HBKU(卡塔尔计算研究所(哈马德·本·哈利法大学))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑