arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过MLLM引导的概念生成和语义传播进行无监督多模态意图发现

Unsupervised Multimodal Intent Discovery via MLLM-Guided Concept Generation and Semantic Propagation

Yunjin Gu, Qianrui Zhou, Hua Xu

arXiv 2607.21908首次发表:更新:

AI 中文总结

该研究针对无监督多模态意图发现缺乏语义监督及可解释性差的问题,提出MCSP方法,通过MLLM引导的对比推理获取语义概念,再经语义传播生成伪标签,实验证明其性能优于现有方法且能产生可解释的簇。

AI 中文摘要

无监督多模态意图发现旨在从无标签的多模态对话中揭示潜在意图,但由于缺乏明确的语义监督而具有挑战性。现有方法的可解释性有限,其优化主要依赖几何相似性而非高级语义指导。为解决这些局限,我们提出MCSP,一种完全无监督的方法,将基于概念的语义优化引入多模态意图发现。为获取可靠的意图发现语义证据,我们为每个簇识别高质量代表性样本,用于支持与相邻簇的MLLM引导的对比推理,产生可解释的高级语义概念。基于这些概念,我们在语义加权图上进行语义传播,使概念信息与局部结构一致性对齐,并生成可靠的伪标签用于表示优化。在三个具有挑战性的多模态意图数据集上的大量实验表明,MCSP始终优于现有方法,同时产生基于语义概念的可解释簇。

英文摘要

Unsupervised multimodal intent discovery aims to uncover latent intents from unlabeled multimodal dialogues, but remains challenging due to the lack of explicit semantic supervision. Existing methods often provide limited interpretability, as their refinement mainly relies on geometric similarity rather than high-level semantic guidance. To address these limitations, we propose MCSP, a fully unsupervised method that introduces semantic refinement based on concepts into multimodal intent discovery. To obtain reliable semantic evidence for intent discovery, we identify high-quality representative samples for each cluster and use them to support MLLM-guided contrastive reasoning against neighboring clusters, which produces interpretable high-level semantic concepts. Building on these concepts, we perform semantic propagation over a semantically weighted graph to align conceptual information with local structural consistency and generate reliable pseudo-labels for representation refinement. Extensive experiments on three challenging multimodal intent datasets show that MCSP consistently outperforms state-of-the-art methods while producing interpretable clusters grounded in semantic concepts.

CommentsAccepted at ACM Multimedia 2026 (MM '26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑