arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-04 至 2025-11-04 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 6 篇

2511.00940 2025-11-04 cs.RO cs.AI 83%

URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model

Zhe Li, Xiang Bai, Jieyu Zhang, Zhuangzhe Wu, Che Xu, Ying Li, Chengkai Hou, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) University of Washington(华盛顿大学)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05335 2025-11-04 cs.CV 74%

New multimodal similarity measure for image registration via modeling local functional dependence with linear combination of learned basis functions

Joel Honkamaa, Pekka Marttinen

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments Improved experimental setup

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15685 2025-11-04 cs.RO 67%

From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systems

Xiuchao Sui, Daiying Tian, Qi Sun, Ruirui Chen, Dongkyu Choi, Kenneth Kwok, Soujanya Poria

机构 * IHPC, Agency for Science, Technology and Research, Singapore(科技研究局智能技术中心,新加坡) Nanyang Technological University, Singapore(南洋理工大学,新加坡)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract)

Comments EMNLP 2025 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00096 2025-11-04 cs.MA cs.AI cs.CY 57%

Urban-MAS: Human-Centered Urban Prediction with LLM-Based Multi-Agent System

Shangyu Lou

机构 * University of California, Santa Barbara \& San Diego State University California USA University of California, Santa Barbara \& San Diego State University

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to The 3rd ACM SIGSPATIAL International Workshop on Advances in Urban AI (UrbanAI'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00933 2025-11-04 cs.RO cs.CV 57%

Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation

Xiangyu Shi, Zerui Li, Yanyuan Qiao, Qi Wu

机构 * Australian Institute for Machine Learning, the University of Adelaide(澳大利亚机器学习研究所、阿德莱德大学) CREATE Lab, Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院CREATE实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00936 2025-11-04 cs.HC 50%

Exploring Human-AI Interaction with Patient-Generated Health Data Sensemaking for Cardiac Risk Reduction

Pavithren V S Pakianathan, Rania Islambouli, Hannah McGowan, Diogo Branco, Tiago Guerreiro, Jan David Smeddinck

专题命中 多模态Agent :multi-modal(abstract)

Comments Presented as demonstration at the workshop on visual analytics in healthcare (VAHC) (in conjunction with IEEE VIS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏