arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-21 至 2025-10-21 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 10 篇

2510.17023 2025-10-21 cs.CV cs.MM 84%

Enrich and Detect: Video Temporal Grounding with Multimodal LLMs

Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa, Lorenzo Torresani, Triantafyllos Afouras

机构 * FAIR, Meta(FAIR、Meta) Johns Hopkins University(约翰霍普金斯大学) Northeastern University(东北大学)

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments ICCV 2025 (Highlights)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17038 2025-10-21 cs.RO cs.AI cs.CV 81%

DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation

Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, Javad Dargahi

机构 * Gina Cody School of Engineering and Computer Science, Concordia University(甘娜·柯迪工程与计算机科学学院,康科迪亚大学) The Wilfred and Joyce Posluns Centre for Image Guided Innovation & Therapeutic Intervention (PCIGITI) at the Hospital for Sick Children (SickKids)(威廉与乔伊斯·波斯卢斯影像引导创新与治疗干预中心(PCIGITI)(SickKids医院)) Electrical and Computer Engineering Department, Kettering University(电气与计算机工程系,凯特林大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16972 2025-10-21 cs.CV cs.AI 73%

The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA

Quanzhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang, Yikang Zhou, Haobo Yuan, Lu Qi, Xiangtai Li, Shunping Ji

机构 * Wuhan University(武汉大学) University of California, Merced(加州大学默塞德分校) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments The 1st place report of 7th LSVOS challenge RVOS track in ICCV 2025. The code is released in Sa2VA repository: https://github.com/bytedance/Sa2VA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16444 2025-10-21 cs.CV cs.MM cs.RO eess.IV 62%

RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba

Kunyu Peng, Di Wen, Jia Fu, Jiamin Wu, Kailun Yang, Junwei Zheng, Ruiping Liu, Yufan Chen, Yuqian Fu, Danda Pani Paudel, Luc Van Gool, Rainer Stiefelhagen

机构 * Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology(人机化研究所,卡尔斯鲁厄技术大学) RISE Research Institutes of Sweden(瑞典RISE研究机构) KTH Royal Institute of Technology(皇家理工学院) School of Artificial Intelligence and Robotics(人工智能与机器人学院) National Engineering Research Center of Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) Chinese University of Hong Kong(香港中文大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Extended version of ECCV 2024 paper arXiv:2407.01872. The dataset and code are released at https://github.com/KPeng9510/refAVA2

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17384 2025-10-21 cs.CV 57%

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

Jiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei Yang

机构 * ShanghaiTech University(上海科技大学) School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17212 2025-10-21 cs.LG cs.AI 57%

D2C-HRHR: Discrete Actions with Double Distributional Critics for High-Risk-High-Return Tasks

Jundong Zhang, Yuhui Situ, Fanji Zhang, Rongji Deng, Tianqi Wei

机构 * School of Artificial Intelligence, Sun Yat-sen University(人工智能学院,中山大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16989 2025-10-21 cs.CV 57%

Training-free Online Video Step Grounding

Luca Zanella, Massimiliano Mancini, Yiming Wang, Alessio Tonioni, Elisa Ricci

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会) Google(谷歌)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025. Project website at https://lucazanella.github.io/baglm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16455 2025-10-21 cs.CL 57%

RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning

Deyi Ji, Yuekui Yang, Haiyang Wu, Shaoping Ma, Tianrun Chen, Lanyun Zhu

机构 * Tencent(腾讯公司) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Zhejiang University(浙江大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments ACL 2025 (Oral, Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16980 2025-10-21 cs.LG 50%

Towards Interpretable and Trustworthy Time Series Reasoning: A BlueSky Vision

Kanghui Ning, Zijie Pan, Yushan Jiang, Anderson Schneider, Yuriy Nevmyvaka, Dongjin Song

机构 * School of Computing University of Connecticut Storrs, CT(计算学院 美国康涅狄格大学 斯托尔斯分校) Department of Machine Learning Research Morgan Stanley New York, NY(机器学习研究部 花旗集团 新 York)

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19949 2025-10-21 eess.IV cs.LG 50%

Automated Video-EEG Analysis in Epilepsy Studies: Advances and Challenges

Valerii A. Zuev, Elena G. Salmagambetova, Stepan N. Djakov, Lev V. Utkin

机构 * Peter the Great St.Petersburg Polytechnic University(彼得大帝圣彼得堡理工大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏