读懂房间,读懂图像:理解多模态视觉语境中的间接言语行为
Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts
浏览论文内容
中文总结 AI 辅助
该研究推出多模态基准READI,以视觉语用问答任务评估间接言语行为理解,发现先进多模态模型在视觉接地间接言语行为上表现差,性能随间接性提升而下降,凸显相关基准的必要性。
中文摘要 AI 辅助
间接言语行为(ISAs)需要对语境进行语用推理,因为指令意图无法仅从字面形式推断。先前基于文本的研究和现有多模态基准大多忽视了这一需求,转而关注明确编码的语境或感知识别,因此未充分探索依赖语境的语用理解,尤其是在韩语等高语境语言中。我们推出READI,这是一个多模态基准,用于通过对视觉语境和对话的综合推理来评估ISA理解。READI基于语用理论建模分级间接性,并将该任务表述为基于视觉的语用问答(V-PQA),支持英语和韩语的跨语言评估。实验表明,即使是最先进的多模态模型在视觉接地的间接言语行为上也表现不佳,且性能随间接性程度增加而下降,凸显了对明确针对语境语用推理的基准的需求。
英文摘要
Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks largely overlook this requirement, focusing instead on explicitly encoded context or perceptual recognition, and thus underex- plore context-dependent pragmatic understand- ing, particularly in high-context languages such as Korean. We introduce READI, a multimodal benchmark for evaluating ISA understanding through integrated reasoning over visual con- text and dialogue. READI models graded in- directness grounded in pragmatic theory and formulates the task as vision-based pragmatic question answering (V-PQA), supporting cross- lingual evaluation in English and Korean. Ex- periments show that even state-of-the-art multi- modal models struggle with visually grounded indirect speech acts, with performance declin- ing as indirectness increases, underscoring the need for benchmarks that explicitly target con- textual pragmatic reasoning.
发表机构
- Yonsei University(延世大学)
- LG AI Research(LG人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。