arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-03 至 2025-12-03 共收录 50 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 5 篇

2512.02438 2025-12-03 cs.CV cs.AI 62%

Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources

通过有限计算资源下的动量自蒸馏提升医学视觉-语言预训练

Phuc Pham, Nhu Pham, Ngoc Quoc Ly

机构 * Faculty of Information Technology, University of Science, Ho Chi Minh City, Vietnam(信息科技学院,科学大学,胡志明市,越南) Vietnam National University, Ho Chi Minh City, Vietnam(越南国家大学,胡志明市,越南)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究通过动量自蒸馏与梯度累积提升医学视觉-语言预训练效率,实现高效多模态学习并提升零样本分类和少量样本适应性能。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02834 2025-12-03 cs.RO cs.AI 57%

Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach

引导视觉-语言-动作模型作为反探索:一种测试时间缩放方法

Siyuan Yang, Yang Zhang, Haoran He, Ling Pan, Xiu Li, Chenjia Bai, Xuelong Li

机构 * Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院) University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出TACO框架,通过测试时间缩放方法在视觉-语言-动作模型中引入反探索机制,提升推理稳定性和任务成功率。

Comments The first two authors contributed equally. Yang Zhang leads the whole project

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03079 2025-12-03 cs.CV 57%

Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA

偏见超越人口统计:通过反事实视觉问答探测黑盒大视觉-语言模型的决策边界

Zaiying Zhao, Toshihiko Yamasaki

机构 * The University of Tokyo(东京大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文通过反事实VQA基准探测黑盒LVLMs的决策边界,揭示非人口属性对决策的更大影响,并展示人类规范验证示例对提升模型响应一致性和公平性的作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02536 2025-12-03 cs.CV 57%

WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens

WeMMU: 通过噪声查询标记增强视觉-语言模型与扩散模型的桥梁

Jian Yang, Dacheng Yin, Xiaoxuan He, Yong Li, Fengyun Rao, Jing Lyu, Wei Zhai, Yang Cao, Zheng-Jun Zha

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知联合实验室,中国科学技术大学) ZheJiang University(浙江大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 WeMMU通过噪声查询标记和VAE分支,提升视觉-语言模型与扩散模型的连接效率,缓解泛化崩溃问题,实现稳定持续学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02517 2025-12-03 cs.CV 57%

SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts

SkyMoE:一种用于增强遥感解释的视觉-语言基础模型

Jiaqi Liu, Ronghao Fu, Lang Sun, Haoran Liu, Xiao Yang, Weipeng Zhang, Xu Na, Zhuoran Duan, Bo Yang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 SkyMoE通过Mixture-of-Experts架构提升遥感多模态多任务处理能力,实现对不同粒度任务的高效适应与优化。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2512.02306 2025-12-03 cs.AI cs.CL cs.CR cs.CV cs.LG 87%

OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning

OmniGuard:统一的多模态防护机制与刻意推理

Boyu Zhu, Xiaofei Wen, Wenjie Jacky Mo, Tinghui Zhu, Yanan Xie, Peng Qi, Muhao Chen

机构 * Fudan University(复旦大学) University of California, Davis(加州大学戴维斯分校)

专题命中 音频语音多模态 :omni-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 OmniGuard通过统一的多模态防护机制与刻意推理能力,有效提升多模态安全防护的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02759 2025-12-03 eess.AS cs.SD eess.IV 86%

Towards Language-Independent Face-Voice Association with Multimodal Foundation Models

面向语言无关的面容-语音关联的多模态基础模型

Aref Farhadipour, Teodora Vukovic, Volker Dellwo

机构 * Department of Computational Linguistics, University of Zurich(苏黎世大学计算语言学系)

专题命中 音频语音多模态 :multimodal(title);multimodal foundation model(title);cross-modal(abstract);分类 eess.AS

AI总结 本文提出了一种基于多模态基础模型的跨语言面容-语音关联方法,通过ImageBind与LoRA结合,实现了在未见过语言上的有效验证。

Comments This paper presents the system description of the UZH-CL team for the FAME2026 Challenge at ICASSP 2026. Our model achieved second place in the final ranking

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02973 2025-12-03 cs.CV cs.CL cs.CR 81%

Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities

基于上下文的图像攻击:如何通过视觉上下文暴露多模态安全漏洞

Yuan Xiong, Ziqi Miao, Lijun Li, Chen Qian, Jie Li, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Xi’an Jiaotong University(西安交通大学) Renmin University of China(中国人民大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出上下文图像攻击方法,通过多智能体系统和四种可视化策略,有效利用图像的上下文信息对多模态大语言模型进行劫持攻击,实验结果显示攻击效果优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02523 2025-12-03 cs.SD cs.LG 71%

Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation

生成式多模态反馈用于歌唱语音合成评估

Xueyan Li, Yuxin Wang, Mengjie Jiang, Qingzi Zhu, Jiang Zhang, Zoey Kim, Yazhe Niu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Science and Technology of China(中国科学技术大学) Columbia University(哥伦比亚大学) Dalian University of Technology(大连理工大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 音频语音多模态 :multi-modal(title)

AI总结 本文提出生成式多模态反馈框架,通过音频-语言模型生成多维语言和音频反馈,提升歌唱语音合成评估的准确性和可解释性。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21025 2025-12-03 cs.SD cs.AI cs.LG eess.AS 62%

Text-Queried Audio Source Separation via Hierarchical Modeling

通过分层建模实现文本查询的音频源分离

Xinlei Yin, Xiulian Peng, Xue Jiang, Zhiwei Xiong, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) School of Information and Communication Engineering, Communication University of China(中国通信大学信息与通信工程学院) Microsoft Research Asia(微软亚洲研究院)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、eess.AS

AI总结 本文提出HSM-TSS框架,通过分层建模实现文本查询的音频源分离,结合双阶段语义分离与结构保持重建,提升复杂场景下的分离性能与语义一致性。

Comments Accepted by TASLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01234 2025-12-03 cs.HC cs.AI 57%

Proactive Agentic Whiteboards: Enhancing Diagrammatic Learning

主动代理白板:增强图示学习

Suveen Ellawela, Sashenka Gamage, Dinithi Dissanayake

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

AI总结 DrawDash通过实时语音驱动的视觉辅助,主动完成和优化教育图表,以减少教师认知负担并提升图示教学效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02515 2025-12-03 cs.SD 50%

VibOmni: Towards Scalable Bone-conduction Speech Enhancement on Earables

VibOmni: 向耳戴设备的可扩展骨传导语音增强迈进

Lixing He, Yunqi Guo, Haozheng Hou, Zhenyu Yan

机构 * department of information engineering, The Chinese University of Hong Kong(信息工程系,香港中文大学)

专题命中 音频语音多模态 :multi-modal(abstract)

AI总结 VibOmni通过骨传导振动与音频融合,提升耳戴设备在嘈杂环境中的语音质量与噪声抑制能力。

Comments Submitted to TMC

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2512.02651 2025-12-03 cs.HC cs.CV cs.SE 79%

Real-Time Multimodal Data Collection Using Smartwatches and Its Visualization in Education

利用智能手表进行实时多模态数据采集及其在教育中的可视化

Alvaro Becerra, Pablo Villegas, Ruth Cobos

机构 * Universidad Autónoma de Madrid(马德里自治大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出利用智能手表进行实时多模态数据采集及可视化,用于教育场景中的学习分析。

Comments Accepted in Technological Ecosystems for Enhancing Multiculturality (TEEM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02609 2025-12-03 cs.RO cs.CV 79%

SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction

SAM2Grasp:通过提示条件化的时序动作预测解决多模态抓取

Shengkai Wu, Jinrong Yang, Wenqiu Luo, Linfeng Gao, Chaohui Shang, Meiyu Zhi, Mingshan Sun, Fangping Yang, Liangliang Ren, Yong Zhao

机构 * CVTE(中国中车)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 SAM2Grasp通过提示条件化的时序动作预测,解决多模态抓取中的冲突问题,实现高精度抓取性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02584 2025-12-03 cs.MM 70%

Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event Extraction

分步模式引导提示框架与参数高效指令微调用于多模态事件提取

Xiang Yuan, Xinrong Chen, Haochen Li, Hang Yang, Guanyu Wang, Weiping Li, Tong Mo

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.MM

AI总结 本文提出分步模式引导提示框架与参数高效指令微调方法,用于提升多模态事件提取任务的性能。

Comments Accepted by 2025 IEEE International Conference on Multimedia and Expo

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05332 2025-12-03 cs.CV cs.CL 62%

Unleashing Hour-Scale Video Training for Long Video-Language Understanding

释放小时级视频训练以实现长视频-语言理解

Jingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang, Jiang Liu, Yusheng Su, Xiaodong Yu, Hao Chen, Jiebo Luo, Zicheng Liu, Emad Barsoum

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出VideoMarathon数据集和Hour-LLaVA模型,通过小时级视频训练提升长视频-语言理解能力。

Comments NeurIPS 2025, Project page: https://videomarathon.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02729 2025-12-03 cs.RO 50%

RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning

RoboWheel: 从现实世界人类示范数据中提取的机器人学习跨形态数据引擎

Yuhong Zhang, Zihan Gao, Shengpeng Li, Ling-Hao Chen, Kaisheng Liu, Runqing Cheng, Xiao Lin, Junjia Liu, Zhuoheng Li, Jingyi Feng, Ziyan He, Jintian Lin, Zheyan Huang, Zhifang Liu, Haoqian Wang

机构 * Tsinghua University(清华大学) Synapath CUHK(香港中文大学) HKU(香港大学) PolyU

专题命中 视频多模态 :multimodal(abstract)

AI总结 RoboWheel通过高精度HOI重建和强化学习优化,将现实世界人类示范转化为跨形态机器人学习的有效数据,提供轻量级、通用的运动表示,提升机器人学习性能。

Comments 27 Pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02260 2025-12-03 q-bio.QM stat.ML 50%

EcoCast: A Spatio-Temporal Model for Continual Biodiversity and Climate Risk Forecasting

EcoCast:一种用于持续生物多样性和气候风险预测的空间时间模型

Hammed A. Akande, Abdulrauf A. Gidado

专题命中 视频多模态 :multi-modal(abstract)

AI总结 EcoCast通过多源数据和序列Transformer模型,实现持续的生物多样性和气候风险预测,提升非洲生态保护策略的科学依据。

Comments 9 pages, 3 figures, 1 table. Accepted to the NeurIPS 2025 Workshop on Tackling Climate Change with Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2411.05036 2025-12-03 cs.CL 79%

From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models

从词向量到多模态嵌入:大型语言模型的技术、应用与未来方向

Charles Zhang, Benji Peng, Xintian Sun, Qian Niu, Junyu Liu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Yichao Zhang, Xinyuan Song, Cheng Fei, Caitlyn Heqi Yin, Lawrence KQ Yan, Hongyang He, Tianyang Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) Simon Fraser University(西蒙弗雷泽大学) Kyoto University(京都大学) National Taiwan Normal University(台湾师范大学) Purdue University(普渡大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校) Emory University(埃默里大学) Cornell University(康奈尔大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The Hong Kong University of Science(香港科学大学) University of Liverpool(利物浦大学) University of Warwick(沃里克大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 本文综述了从词向量到多模态嵌入的发展,探讨了大型语言模型的技术、应用及未来方向。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2512.02351 2025-12-03 cs.CV cs.AI 81%

Understanding and Harnessing Sparsity in Unified Multimodal Models

理解并利用统一多模态模型中的稀疏性

Shwai He, Chaorui Deng, Ang Li, Shen Yan

机构 * ByteDance Seed(字节跳动种子) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究通过MoE适应方法,利用稀疏激活提升统一多模态模型的效率,使模型在激活约一半参数的情况下达到与完整模型相当的性能。

Comments 13 pages, 13 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02088 2025-12-03 eess.IV cs.AI cs.CV cs.LG 81%

Comparing Baseline and Day-1 Diffusion MRI Using Multimodal Deep Embeddings for Stroke Outcome Prediction

比较基线和第1天扩散磁共振成像用于中风预后预测的多模态深度嵌入

Sina Raeisadigh, Myles Joshua Toledo Tan, Henning Müller, Abderrahmane Hedjoudje

机构 * 1 Department of Computer Science, University of Geneva, Switzerland 2 Department of Electrical \& Computer Engineering, University of Florida, FL, USA 3 Service of Medical Informatics, University Hospital of Geneva, Switzerland 4 Department of Imaging Medical Informatics, University of Geneva, Switzerland

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究通过多模态深度嵌入方法,利用基线和治疗后1天的扩散MRI数据,结合临床特征和病变体积,预测急性缺血性中风患者3个月的功能预后。

Comments 5 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02273 2025-12-03 cs.CV cs.AI 62%

Progressive Image Restoration via Text-Conditioned Video Generation

通过文本条件视频生成实现渐进式图像修复

Peng Kang, Xijun Wang, Yu Yuan

机构 * Department of Computer Science, University of Illinois Springfield(伊利诺伊大学斯普林菲尔德分校计算机科学系) School of Electrical and Computer Engineering, Purdue University(普渡大学电气与计算机工程学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本研究通过微调CogVideo,利用文本条件视频生成实现图像渐进式修复,提升感知度量并展示强泛化能力。

Comments First two authors contributed equally to this work. IEEE ICNC Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06996 2025-12-03 cs.CV cs.AI 62%

Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems

可见却不可读:跨书写系统下视觉语言模型的系统性盲区

Jie Zhang, Ting Xu, Gelei Deng, Runyi Hu, Han Qiu, Tianwei Zhang, Qing Guo, Ivor Tsang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了视觉语言模型在跨书写系统下识别碎片化文本的鲁棒性,发现其在可见但不可读的刺激下表现下降,揭示了模型对组成先验依赖不足的结构性限制。

Comments arXiv admin note: This article has been withdrawn by arXiv administrators due to violation of arXiv policy regarding generative AI authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02713 2025-12-03 cs.AI 57%

Training Data Attribution for Image Generation using Ontology-Aligned Knowledge Graphs

利用本体对齐的知识图谱训练数据归因于图像生成

Theodoros Aivalis, Iraklis A. Klampanos, Antonis Troumpoukis, Joemon M. Jose

机构 * National Centre for Scientific Research ``Demokritos''(国家科学研究中心「德莫克里特」) University of Glasgow(格拉斯哥大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出利用本体对齐的知识图谱方法,通过多模态大语言模型提取图像中的结构化三元组,以追踪生成模型中训练数据的影响,从而提升透明度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02554 2025-12-03 cs.CV 57%

OmniPerson: Unified Identity-Preserving Pedestrian Generation

OmniPerson: 统一的身份保持行人生成

Changxiao Ma, Chao Yuan, Xincheng Shi, Yuzhuo Ma, Yongfei Zhang, Longkun Zhou, Yujia Zhang, Shangze Li, Yifan Xu

机构 * Beihang University(北航大学) Pengcheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 OmniPerson通过统一身份保持的行人生成管道,解决ReID中数据不足问题,实现高保真生成和身份一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02492 2025-12-03 cs.CV 57%

YingVideo-MV: Music-Driven Multi-Stage Video Generation

YingVideo-MV: 基于音乐的多阶段视频生成

Jiahui Chen, Weida Wang, Runhua Shi, Huan Yang, Chaofan Ding, Zihao Chen

机构 * AI Lab, GiantNetwork(人工智能实验室)

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV

AI总结 YingVideo-MV通过整合音频语义分析、时间感知扩散模型和摄像机运动控制模块,实现了基于音乐的高质量视频生成。

Comments 18 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08096 2025-12-03 cs.CV 57%

MegaSR: Mining Customized Semantics and Expressive Guidance for Real-World Image Super-Resolution

MegaSR: 为真实世界图像超分辨率挖掘定制语义和表达引导

Xinrui Li, Jinrong Zhang, Jianlong Wu, Chong Chen, Liqiang Nie, Zhouchen Lin

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 MegaSR通过定制语义和表达引导提升真实世界图像超分辨率的重建质量与结构一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 10 篇

2512.02668 2025-12-03 cs.CV 83%

UAUTrack: Towards Unified Multimodal Anti-UAV Visual Tracking

UAUTrack:迈向统一的多模态反无人机视觉跟踪

Qionglin Ren, Dawei Zhang, Chunxu Tian, Dan Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) Fudan University(复旦大学) School of Computer Science and Technology(计算机科学与技术学院) Zhejiang Normal University(浙江师范大学) Department of Mechanical Engineering(机械工程系) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 UAUTrack通过统一多模态框架实现反无人机跟踪,结合文本先验提示策略提升跨模态信息整合效率,取得优异性能与效率平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19551 2025-12-03 cs.CY cs.AI cs.CV 81%

Rainbow Noise: Stress-Testing Multimodal Harmful-Meme Detectors on LGBTQ Content

彩虹噪声:对多模态有害迷因检测器进行压力测试以应对LGBTQ内容

Ran Tong, Songtao Wei, Jiaqi Liu, Lanruo Wang

机构 * Mathematics and Statistics Department University of Texas at Dallas(德克萨斯大学达拉斯分校数学与统计学系) Computer Science Department University of Texas at Dallas(德克萨斯大学达拉斯分校计算机科学系) Naveen Jindal School of Management University of Texas at Dallas(德克萨斯大学达拉斯分校奈文·金达尔管理学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出彩虹噪声基准测试,通过压力测试多模态有害迷因检测器,引入轻量级TDA提升鲁棒性,揭示当前模型弱点并展示改进方向。

Comments 14 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02699 2025-12-03 cs.AI 79%

Learning What to Attend First: Modality-Importance-Guided Reasoning for Reliable Multimodal Emotion Understanding

学习优先关注什么:模态重要性引导的推理用于可靠的多模态情感理解

Hyeongseop Rha, Jeong Hun Yeo, Junil Won, Se Jin Park, Yong Man Ro

机构 * Integrated Vision and Language Lab, KAIST, South Korea(韩国成均馆大学集成视觉与语言实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出MIGR框架,通过模态重要性引导推理,提升多模态情感理解的可靠性,实验表明能有效减少情感不一致的解释实例。

Comments 16 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏