arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-21 至 2025-11-21 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 6 篇

2511.16635 2025-11-21 cs.CV cs.CL 84%

SurvAgent: Hierarchical CoT-Enhanced Case Banking and Dichotomy-Based Multi-Agent System for Multimodal Survival Prediction

SurvAgent: 基于层次化CoT增强的案例库与二元多智能体系统用于多模态生存预测

Guolin Huang, Wenting Chen, Jiaqi Yang, Xinheng Lyu, Xiaoling Luo, Sen Yang, Xiaohan Xing, Linlin Shen

机构 * Shenzhen University(深圳大学) Stanford University(斯坦福大学) University of Nottingham Ningbo China(诺丁汉大学宁波分校) Ant Group(蚂蚁集团)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 SurvAgent通过层次化CoT增强的多智能体系统,整合多模态数据,提升生存预测的可解释性与准确性。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16183 2025-11-21 cs.AI cs.CV 81%

FOOTPASS: A Multi-Modal Multi-Agent Tactical Context Dataset for Play-by-Play Action Spotting in Soccer Broadcast Videos

FOOTPASS:一个多模态多智能体战术上下文数据集,用于足球比赛视频中的 play-by-play 动作识别

Jeremie Ochin, Raphael Chekroun, Bogdan Stanciulescu, Sotiris Manitsaris

机构 * Center for Robotics, Mines Paris - PSL(机器人中心,巴黎-PSL Mines)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 FOOTPASS数据集旨在通过多模态和多智能体战术上下文,实现足球比赛视频中play-by-play动作的自动化识别与可靠提取。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17425 2025-11-21 cs.AI 79%

Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments

在家庭环境中评估多模态大语言模型的日常复合任务

Zhenliang Zhang, Yuxi Wang, Hongzhao Xie, Shiyun Zhao, Mingyuan Liu, Yujie Lu, Xinyi He, Zhenku Cheng, Yujia Peng

机构 * State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence(通用人工智能国家重点实验室、北京通用人工智能研究院) School of Psychological and Cognitive Sciences and Beijing Key Laboratory of Behavior and Mental Health, Key Laboratory of Machine Perception (Ministry of Education), Peking University(心理与认知科学学院及北京行为与心理健康重点实验室、机器感知重点实验室(教育部)) School of Intelligence Science and Technology, Peking University(智能科学与技术学院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本研究在家庭环境中评估了多模态大语言模型在日常复合任务中的表现,发现其在物体理解、空间智能和社会活动领域存在显著差距,为具身MLLMs的发展提供了初步评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14210 2025-11-21 cs.CV cs.AI cs.LG 76%

Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution

Orion:一个多模态感知、高级视觉推理与执行的统一视觉代理

N Dinesh Reddy, Dylan Snyder, Lona Kiragu, Mirajul Mohin, Shahrear Bin Amin, Sudeep Pillai

机构 * VLM Run Research(VLM Run研究)

专题命中 多模态Agent :multimodal(title);分类 cs.CV、cs.AI

AI总结 Orion通过整合视觉推理与工具增强执行,实现了多模态感知、高级视觉推理与执行的统一,提升多步骤视觉智能性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15997 2025-11-21 cs.AI cs.MM 62%

Sensorium Arc: AI Agent System for Oceanic Data Exploration and Interactive Eco-Art

Sensorium Arc:面向海洋数据探索与交互生态艺术的AI代理系统

Noah Bissell, Ethan Paley, Joshua Harrison, Juliano Calil, Myungin Lee

机构 * Immersive Media Design University of Maryland College Park(沉浸媒体设计大学马里兰大学学院公园分校) Center for the Study of the Force Majeure University of California, Santa Cruz(重大研究机构加州大学圣克鲁兹分校) Virtual Planet Technologies(虚拟星球技术)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI、cs.MM

AI总结 Sensorium Arc通过AI代理将海洋数据转化为生动叙事,结合科学洞察与生态诗意,实现沉浸式环境数据探索与交互生态艺术

Comments (to appear) NeurIPS 2025 Creative AI Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16063 2025-11-21 cs.NI 50%

Modeling Pointing, Acquisition, and Tracking Delays in Free-Space Optical Satellite Networks

自由空间光学卫星网络中指针、获取和跟踪延迟的建模

Jason Gerard, Juan A. Fraire, Sandra Céspedes

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种用于自由空间光学卫星网络中指针、获取和跟踪延迟建模的验证模型,以提高接触计划的准确性和网络利用率。

Comments 2025 IEEE International Conference on Wireless for Space and Extreme Environments (WiSEE) - STINT Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏