arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-17 至 2026-02-17 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 13 篇

2602.10551 2026-02-17 cs.CV cs.AI 81%

C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning

C^2ROPE: 3D 大多模态模型推理中的因果连续旋转位置编码

Guanting Ye, Qiyan Zhao, Wenhao Yu, Xiaofeng Zhang, Jianmin Ji, Yanyong Zhang, Ka-Veng Yuen

机构 * State Key Laboratory of Internet of Things for Smart City, University of Macau(物联网智能城市国家重点实验室,澳门大学) Department of Automation, Shanghai Jiaotong University(上海交通大学自动化系) Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院) School of Computer Science and Technology, USTC(中国科学技术大学计算机科学与技术学院) School of Artificial Intelligence and Data Science, USTC(中国科学技术大学人工智能与数据科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 C^2ROPE通过引入空间-时间连续位置编码和切比雪夫因果掩码,解决3D多模态模型中视觉特征连续性和因果关系建模问题。

Comments Accepted in ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04641 2026-02-17 cs.CV cs.AI cs.LG 81%

Simulating the Real World: A Unified Survey of Multimodal Generative Models

模拟现实世界:多模态生成模型的统一综述

Yuqi Hu, Longguang Wang, Xian Liu, Ling-Hao Chen, Yuwei Guo, Yukai Shi, Ce Liu, Anyi Rao, Zeyu Wang, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能前沿技术研究所,香港科学与技术大学(广州)) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR(计算机科学与工程系,香港科学与技术大学香港特别行政区) MMLab, The Hong Kong University of Science and Technology(多模态实验室,香港科学与技术大学) School of Electronics and Communication Engineering, Shenzhen Campus of Sun Yat-sen University(电子与通信工程学院,中山大学深圳校区) The Chinese University of Hong Kong, Hong Kong, China(香港中文大学,香港,中国) Tsinghua University, Guangdong, China(清华大学,广东,中国) Bosch (China) Investment Co., Ltd., Shanghai, China(博世(中国)投资有限公司,上海,中国)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文首次系统性地统一研究了2D、视频、3D和4D生成,为多模态生成模型和现实世界模拟提供了统一框架的综述。

Comments Repository for the related papers at https://github.com/ALEEEHU/World-Simulator

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14589 2026-02-17 cs.AI cs.CL cs.LG 81%

MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs

MATEO:一种多模态基准,用于LVLMs中的时间推理和规划

Gabriel Roccabruna, Olha Khomyn, Giuseppe Riccardi

机构 * Signals and Interactive Systems Lab, University of Trento, Italy(特伦托大学信号与交互系统实验室) University of Trento(特伦托大学) Amazon(亚马逊)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 MATEO是一个多模态基准,用于评估和提升大型视觉语言模型在时间推理和规划方面的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21842 2026-02-17 cs.CV cs.CR 79%

Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?

模态失语:统一多模态模型能否从记忆中描述图像?

Michael Aerni, Joshua Swanson, Kristina Nikolić, Florian Tramèr

机构 * Michael Aerni(独立研究者) Joshua Swanson(独立研究者) Kristina Nikolić(独立研究者) Florian Tramèr(独立研究者)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 研究发现统一多模态模型在视觉记忆与文本表达间存在系统性缺陷,导致安全框架可能因单一模态防护而失效。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.12433 2026-02-17 cs.RO cs.SY eess.SY 78%

Model Predictive Control with Gaussian Processes for Flexible Multi-Modal Physical Human Robot Interaction

基于高斯过程的模型预测控制用于柔性多模态人机协作交互

Kevin Haninger, Christian Hegeler, Luka Peternel

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 本文提出基于高斯过程的模型预测控制方法,用于多模态人机协作交互,通过贝叶斯推断和在线控制提升任务灵活性和效率。

Comments Submitted, ICRA 2022. Video: https://youtu.be/0GT1pPpXvt8 Data and code: https://owncloud.fraunhofer.de/index.php/s/kmCZvlKOghclHy9

Journal ref 2022 IEEE International Conference on Robotics and Automation (ICRA), May 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13459 2026-02-17 eess.SP 78%

Towards Causality-Aware Modeling for Multimodal Brain-Muscle Interactions

迈向因果意识的多模态脑-肌相互作用建模

Farwa Abbas, Wei Dai, Zoran Cvetkovic, Verity McClelland

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 本文提出一种结合几何流形重建与概率时间建模的DBN启发CCM框架,用于多模态脑-肌交互的因果建模,揭示了肌张力障碍中特定频率的通路重组织,并展示了其在生物标志物开发和神经调节干预中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13329 2026-02-17 cs.CV cs.AI cs.RO 62%

HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving

HiST-VLA:一种用于端到端自动驾驶的分层时空视觉-语言-动作模型

Yiru Wang, Zichong Gu, Yu Gao, Anqing Jiang, Zhigang Sun, Shuo Wang, Yuwen Heng, Hao Sun

机构 * Bosch Corporate Research(博世企业研究) School of Communication and Information Engineering(信息工程学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 HiST-VLA通过分层时空视觉-语言-动作模型提升自动驾驶轨迹生成的精度与效率,实现端到端的自动驾驶系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23232 2026-02-17 cs.CV cs.AI 62%

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

ShotFinder: 通过网络搜索驱动的开放域视频镜头检索

Tao Yu, Haopeng Jin, Hao Wang, Shenghua Chai, Yujia Yang, Junhao Gong, Jiaming Guo, Minghui Zhang, Xinlong Chen, Zhenghao Zhang, Yuxuan Zhou, Yufei Xiong, Shanbin Zhang, Jiabing Yang, Hongzhu Yi, Xinming Wang, Cheng Zhong, Xiao Ma, Zhang Zhang, Yan Huang, Liang Wang

机构 * CASIA(中国科学院自动化研究所) UCAS(中国科学院大学) Lenovo(联想集团) Peking University(北京大学) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ShotFinder通过网络搜索驱动开放域视频镜头检索,提出三阶段检索流程并揭示多模态大模型在开放域视频检索中的性能差距。

Comments 28 pages, 7 figures, Project website: https://github.com/yutao1024/ShotFinder

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14653 2026-02-17 cs.CL 57%

Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?

当话语基于感知和话语进行 grounding 时,信息密度是否均匀?

Matteo Gay, Coleman Haley, Mario Giulianelli, Edoardo Ponti

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

AI总结 本研究首次探讨了基于感知和话语的视觉环境对信息密度均匀性的影响,发现 grounding 能提高信息分布的均匀性,并在话语单元开始处产生最大的惊奇度降低。

Comments Accepted as main paper at EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14214 2026-02-17 cs.CV 57%

HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming

HiVid: 基于大语言模型的视频显著性识别用于内容感知点播与直播

Jiahui Chen, Bo Peng, Lianchen Jia, Zeyu Zhang, Tianchi Huang, Lifeng Sun

机构 * Tsinghua University(清华大学) The Australian National University(澳大利亚国立大学) Key Laboratory of Pervasive Computing, Ministry of Education(教育部普适计算重点实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 HiVid利用大语言模型生成内容感知的视频显著性权重,提升点播和直播的QoE性能。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13496 2026-02-17 cs.CY cs.AI 57%

Future of Edge AI in biodiversity monitoring

边缘AI在生物多样性监测中的未来

Aude Vuilliomenet, Kate E. Jones, Duncan Wilson

机构 * The Bartlett Centre for Advanced Spatial Analysis, Faculty of the Built Environment, University College London(大学学院) Centre for Biodiversity and Environment Research, Department of Genetics, Evolution and Environment, University College London(大学学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文研究了边缘AI在生物多样性监测中的应用,分析了不同系统类型及其权衡,强调了跨学科合作的重要性。

Comments 41 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14255 2026-02-17 cs.RO 50%

A Latency-Aware Framework for Visuomotor Policy Learning on Industrial Robots

面向工业机器人视觉-运动策略学习的延迟感知框架

Daniel Ruan, Salma Mozaffari, Sigrid Adriaenssens, Arash Adel

机构 * Princeton University(普林斯顿大学)

专题命中 视频多模态 :multimodal(abstract)

AI总结 本文提出了一种面向工业机器人视觉-运动策略学习的延迟感知框架,通过优化执行策略以应对延迟问题,提升策略在现实环境中的可靠性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02456 2026-02-17 cs.RO 50%

InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation

InternVLA-A1:统一理解、生成和行动以实现机器人操作

Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, Yanan Lu, Qi Lv, Haoxiang Ma, Jiangmiao Pang, Yu Qiao, Zherui Qiu, Yanqing Shen, Xu Shi, Yang Tian, Bolun Wang, Hanqing Wang, Jiaheng Wang, Tai Wang, Xueyuan Wei, Chao Wu, Yiman Xie, Boyang Xing, Yuqiang Yang, Yuyin Yang, Qiaojun Yu, Feng Yuan, Jia Zeng, Jingjing Zhang, Shenghan Zhang, Shi Zhang, Zhuoma Zhaxi, Bowen Zhou, Yuanzhen Zhou, Yunsong Zhou, Hongrui Zhu, Yangkun Zhu, Yuchen Zhu

机构 * InternVLA-A1 Team(InternVLA-A1团队)

专题命中 视频多模态 :multimodal(abstract)

AI总结 InternVLA-A1通过统一的Transformer架构,结合语义理解和动态预测能力,提升了机器人操作任务的性能。

Comments Homepage: https://internrobotics.github.io/internvla-a1.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏