arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-21 至 2025-11-21 共收录 45 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9 篇

2511.16031 2025-11-21 cs.CV 57%

Crossmodal learning for Crop Canopy Trait Estimation

跨模态学习用于作物冠层特性估计

Timilehin T. Ayanlade, Anirudha Powadi, Talukder Z. Jubery, Baskar Ganapathysubramanian, Soumik Sarkar

机构 * Department of Computer Engineering, Iowa State University, Ames, IA, USA(计算机工程系,爱荷华州立大学) Department of Mechanical Engineering, Iowa State University, Ames, IA, USA(机械工程系,爱荷华州立大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出跨模态学习方法,通过融合高分辨率卫星影像与无人机影像细节,提升作物冠层特性估计的准确性。

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15847 2025-11-21 cs.LG 50%

Transparent Early ICU Mortality Prediction with Clinical Transformer and Per-Case Modality Attribution

透明的早期ICU死亡预测:结合临床Transformer和病例级模态归因

Alexander Bakumenko, Janine Hoelscher, Hudson Smith

机构 * Clemson University(克莱姆森大学)

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出一种透明的多模态集成模型,结合临床Transformer和病例级模态归因,提升ICU早期死亡预测的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 6 篇

2511.16635 2025-11-21 cs.CV cs.CL 84%

SurvAgent: Hierarchical CoT-Enhanced Case Banking and Dichotomy-Based Multi-Agent System for Multimodal Survival Prediction

SurvAgent: 基于层次化CoT增强的案例库与二元多智能体系统用于多模态生存预测

Guolin Huang, Wenting Chen, Jiaqi Yang, Xinheng Lyu, Xiaoling Luo, Sen Yang, Xiaohan Xing, Linlin Shen

机构 * Shenzhen University(深圳大学) Stanford University(斯坦福大学) University of Nottingham Ningbo China(诺丁汉大学宁波分校) Ant Group(蚂蚁集团)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 SurvAgent通过层次化CoT增强的多智能体系统,整合多模态数据,提升生存预测的可解释性与准确性。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16183 2025-11-21 cs.AI cs.CV 81%

FOOTPASS: A Multi-Modal Multi-Agent Tactical Context Dataset for Play-by-Play Action Spotting in Soccer Broadcast Videos

FOOTPASS:一个多模态多智能体战术上下文数据集,用于足球比赛视频中的 play-by-play 动作识别

Jeremie Ochin, Raphael Chekroun, Bogdan Stanciulescu, Sotiris Manitsaris

机构 * Center for Robotics, Mines Paris - PSL(机器人中心,巴黎-PSL Mines)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 FOOTPASS数据集旨在通过多模态和多智能体战术上下文,实现足球比赛视频中play-by-play动作的自动化识别与可靠提取。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17425 2025-11-21 cs.AI 79%

Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments

在家庭环境中评估多模态大语言模型的日常复合任务

Zhenliang Zhang, Yuxi Wang, Hongzhao Xie, Shiyun Zhao, Mingyuan Liu, Yujie Lu, Xinyi He, Zhenku Cheng, Yujia Peng

机构 * State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence(通用人工智能国家重点实验室、北京通用人工智能研究院) School of Psychological and Cognitive Sciences and Beijing Key Laboratory of Behavior and Mental Health, Key Laboratory of Machine Perception (Ministry of Education), Peking University(心理与认知科学学院及北京行为与心理健康重点实验室、机器感知重点实验室(教育部)) School of Intelligence Science and Technology, Peking University(智能科学与技术学院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本研究在家庭环境中评估了多模态大语言模型在日常复合任务中的表现,发现其在物体理解、空间智能和社会活动领域存在显著差距,为具身MLLMs的发展提供了初步评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14210 2025-11-21 cs.CV cs.AI cs.LG 76%

Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution

Orion:一个多模态感知、高级视觉推理与执行的统一视觉代理

N Dinesh Reddy, Dylan Snyder, Lona Kiragu, Mirajul Mohin, Shahrear Bin Amin, Sudeep Pillai

机构 * VLM Run Research(VLM Run研究)

专题命中 多模态Agent :multimodal(title);分类 cs.CV、cs.AI

AI总结 Orion通过整合视觉推理与工具增强执行,实现了多模态感知、高级视觉推理与执行的统一,提升多步骤视觉智能性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15997 2025-11-21 cs.AI cs.MM 62%

Sensorium Arc: AI Agent System for Oceanic Data Exploration and Interactive Eco-Art

Sensorium Arc:面向海洋数据探索与交互生态艺术的AI代理系统

Noah Bissell, Ethan Paley, Joshua Harrison, Juliano Calil, Myungin Lee

机构 * Immersive Media Design University of Maryland College Park(沉浸媒体设计大学马里兰大学学院公园分校) Center for the Study of the Force Majeure University of California, Santa Cruz(重大研究机构加州大学圣克鲁兹分校) Virtual Planet Technologies(虚拟星球技术)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI、cs.MM

AI总结 Sensorium Arc通过AI代理将海洋数据转化为生动叙事,结合科学洞察与生态诗意,实现沉浸式环境数据探索与交互生态艺术

Comments (to appear) NeurIPS 2025 Creative AI Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16063 2025-11-21 cs.NI 50%

Modeling Pointing, Acquisition, and Tracking Delays in Free-Space Optical Satellite Networks

自由空间光学卫星网络中指针、获取和跟踪延迟的建模

Jason Gerard, Juan A. Fraire, Sandra Céspedes

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种用于自由空间光学卫星网络中指针、获取和跟踪延迟建模的验证模型,以提高接触计划的准确性和网络利用率。

Comments 2025 IEEE International Conference on Wireless for Space and Extreme Environments (WiSEE) - STINT Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 6 篇

2511.16229 2025-11-21 cs.CR cs.AI 89%

Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security

Q-MLLM:向量量化用于鲁棒多模态大语言模型安全

Wei Zhao, Zhe Li, Yige Li, Jun Sun

机构 * Singapore Management University(新加坡管理大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 Q-MLLM通过两级向量量化提升多模态大语言模型的安全性,有效防御对抗性攻击并保持模型实用性。

Comments Accepted by NDSS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15741 2025-11-21 cs.AI cs.HC cs.LG 88%

Uncertainty-Resilient Multimodal Learning via Consistency-Guided Cross-Modal Transfer

通过一致性引导的跨模态转移实现不确定性鲁棒的多模态学习

Hyo-Jeong Jang

机构 * Brain and Cognitive Engineering(脑科学与认知工程)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 本论文提出通过一致性引导的跨模态转移实现不确定性鲁棒的多模态学习,旨在提高模型的稳定性和鲁棒性。

Comments Master's thesis, Korea University, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11706 2025-11-21 cs.LG cs.CV 74%

Context-Aware Multimodal Representation Learning for Spatio-Temporally Explicit Environmental Modelling

面向情境的多模态表示学习用于时空明确的环境建模

Julia Peters, Karin Mora, Miguel D. Mahecha, Chaonan Ji, David Montero, Clemens Mosig, Guido Kraemer

机构 * Environmental Data Science and Remote Sensing Group(环境数据科学与遥感小组) Institute for Earth System Science and Remote Sensing(地球系统科学与遥感研究所) Leipzig University(莱比锡大学) German Centre for Integrative Biodiversity Research (iDiv)(整合生物多样性研究德国中心(iDiv))

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 本文提出一种面向情境的多模态表示学习框架,整合不同地球观测模态以高时空分辨率建模环境,提升生态分析的精度与效率。

Comments 10 pages (incliding 2 pages of references), 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16435 2025-11-21 cs.CV 70%

Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation

超越视觉线索:利用通用语义作为少样本分割的支持

Jin Wang, Bingfeng Zhang, Jian Pang, Mengyu Liu, Honglong Chen, Weifeng Liu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出语言驱动属性泛化架构,通过多属性增强和多模态对齐提升少样本分割性能,实现新的最佳效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15967 2025-11-21 cs.CV 57%

InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer

InfoCLIP: 通过信息论对齐转移连接视觉语言预训练与开放词汇语义分割

Muyao Yuan, Yuanhong Zhang, Weizhan Zhang, Lan Ma, Yuan Gao, Jiangyong Ying, Yudeng Xin

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

AI总结 InfoCLIP通过信息论对齐转移提升开放词汇语义分割的性能,有效解决预训练CLIP在微调过程中的过拟合问题。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15669 2025-11-21 cs.CV 57%

UINO-FSS: Unifying Representation Learning and Few-shot Segmentation via Hierarchical Distillation and Mamba-HyperCorrelation

UINO-FSS: 通过层次化蒸馏和Mamba-超相关性统一表示学习与少样本分割

Wei Zhuo, Zhiyue Tang, Wufeng Xue, Hao Ding, Junkai Ji, Linlin Shen

机构 * School of Artificial Intelligence and the National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(人工智能学院和大数据系统计算技术国家工程实验室,深圳大学) Guangdong Provincial Key Laboratory of Intelligent Information Processing, China(广东省智能信息处理重点实验室,中国) School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University(生物医学工程学院,深圳大学医学院,深圳大学) Department of Computer Science, University of Nottingham Ningbo China(计算机科学系,宁波大学中国)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 UINO-FSS通过层次化蒸馏和Mamba-超相关性整合不同基础模型知识,实现少样本分割的统一学习框架,取得新SOTA结果。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 1 篇

2310.16018 2025-11-21 cond-mat.str-el 50%

Symmetry-breaking pathway towards the unpinned broken helix

打破对称性的路径通往未固定的断裂螺旋

E. Donoway, T. V. Trevisan, A. Liebman - Peláez, R. P. Day, K. Yamakawa, Y. Sun, J. R. Soh, D. Prabhakaran, A. T. Boothroyd, R. M. Fernandes, J. G. Analytis, J. E. Moore, J. Orenstein, V. Sunko

专题命中 其他多模态 :multimodal(abstract)

AI总结 研究通过多模式方法揭示EuIn₂As₂中高温和低温磁结构的对称性破缺,发现未固定的断裂螺旋态可恢复轴子相。

Comments 32 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏