arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-21 至 2025-11-21 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9 篇

2511.16221 2025-11-21 cs.CV cs.CL 81%

Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions

大语言模型能读取环境吗?一个多模态基准用于评估多当事人社交互动中的欺骗

Caixin Kang, Yifei Huang, Liangyang Ouyang, Mingfang Zhang, Ruicong Liu, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出多模态互动欺骗评估任务,通过新颖数据集评估多种MLLMs的欺骗检测能力,揭示其在多模态社交线索处理上的不足,并提出SoCoT和DSEM模块提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19390 2025-11-21 eess.SP 78%

Depression diagnosis from patient interviews using multimodal machine learning

基于多模态机器学习的抑郁症诊断

Jana Weber, Marcel Weber, Juan Miguel Lopez Alcaraz

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本研究通过多模态机器学习整合语音、语言和临床信息,提升抑郁症诊断的准确性与临床实用性。

Comments Accepted by Frontiers in Psychiatry, 19 pages, 5 figures, source code under https://github.com/UOLMDA2025/Depression

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15205 2025-11-21 cs.LG cs.AI cs.CL 62%

Long-Short Distance Graph Neural Networks and Improved Curriculum Learning for Emotion Recognition in Conversation

长短期距离图神经网络与改进的课程学习用于对话中的情绪识别

Xinran Li, Xiujuan Xu, Jiaqi Qiao

机构 * School of Software Technology, Dalian University of Technology(大连理工大学软件学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出长短期距离图神经网络与改进课程学习方法,用于对话中情绪识别,通过多模态特征提取和优化训练策略提升模型性能。

Comments Accepted by the 28th European Conference on Artificial Intelligence (ECAI 2025)

Journal ref ECAI 2025, Frontiers in Artificial Intelligence and Applications, Volume 413, pp. 4033-4040, IOS Press, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16440 2025-11-21 cs.CV 57%

StreetView-Waste: A Multi-Task Dataset for Urban Waste Management

StreetView-Waste: 一个用于城市垃圾管理的多任务数据集

Diogo J. Paulo, João Martins, Hugo Proença, João C. Neves

机构 * University of Beira Interior(贝拉内陆大学) IT: Instituto de Telecomunicações(电信研究所) NOVA LINCS

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 StreetView-Waste数据集通过多任务设置,提供垃圾桶检测、跟踪和溢出分割的基准,结合启发式方法和几何先验提升模型性能。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14945 2025-11-21 cs.CV 57%

Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities

无监督发现人类活动中的长期时空周期性工作流

Fan Yang, Quanting Xie, Atsunori Moteki, Shoichi Masui, Shan Jiang, Kanji Uchino, Yonatan Bisk, Graham Neubig

机构 * Fujitsu Research of America, USA(美国富士通研究机构) Fujitsu Limited, Japan(日本富士通有限公司) Carnegie Mellon University, USA(美国卡内基梅隆大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出首个包含长期周期性工作流的基准,通过轻量级基线实现无监督检测和异常检测,显著优于现有方法并具备实际部署优势。

Comments accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16163 2025-11-21 cs.CV 57%

An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs

一张图像等于一万字:针对视觉语言模型的冗余文本诱导攻击

Zhi Luo, Zenghui Yuan, Wenqi Wei, Daizong Liu, Pan Zhou

机构 * Huazhong University of Science and Technology(华中科技大学) Fordham University(福特汉姆大学) Wuhan University(武汉大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种针对视觉语言模型的冗余文本诱导攻击,通过两阶段框架生成恶意图像以诱导模型生成冗长文本,提升攻击效果与可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16048 2025-11-21 cs.RO cs.AI cs.HC 57%

Semantic Glitch: Agency and Artistry in an Autonomous Pixel Cloud

语义裂隙:自主像素云中的代理与艺术性

Qing Zhang, Jing Huang, Mingyang Xu, Jun Rekimoto

机构 * The University of Tokyo(东京大学) Tokyo University of the Arts(东京艺术大学) Keio University(庆应大学) SONY CSL Kyoto(索尼 CSL 京都)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 本文提出语义裂隙,通过多模态大语言模型实现自主导航,展示低 fidelity方法在创造不完美但富有艺术性的机器人同伴中的应用。

Comments NeurIPS 2025 Creative AI Track, The Thirty-Ninth Annual Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16031 2025-11-21 cs.CV 57%

Crossmodal learning for Crop Canopy Trait Estimation

跨模态学习用于作物冠层特性估计

Timilehin T. Ayanlade, Anirudha Powadi, Talukder Z. Jubery, Baskar Ganapathysubramanian, Soumik Sarkar

机构 * Department of Computer Engineering, Iowa State University, Ames, IA, USA(计算机工程系,爱荷华州立大学) Department of Mechanical Engineering, Iowa State University, Ames, IA, USA(机械工程系,爱荷华州立大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出跨模态学习方法,通过融合高分辨率卫星影像与无人机影像细节,提升作物冠层特性估计的准确性。

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15847 2025-11-21 cs.LG 50%

Transparent Early ICU Mortality Prediction with Clinical Transformer and Per-Case Modality Attribution

透明的早期ICU死亡预测:结合临床Transformer和病例级模态归因

Alexander Bakumenko, Janine Hoelscher, Hudson Smith

机构 * Clemson University(克莱姆森大学)

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出一种透明的多模态集成模型,结合临床Transformer和病例级模态归因,提升ICU早期死亡预测的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏