arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-03 至 2026-02-03 共收录 153 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 18 篇

2602.00559 2026-02-03 cs.CV cs.AI 81%

Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models

在视频多模态大语言模型中学习对抗组合性幻觉

Wenbin Xing, Quanxing Zha, Lizheng Zu, Mengran Li, Ming Li, Junchi Yan

机构 * Sun Yat-sen University(中山大学) Huaqiao University(华侨大学) Shenzhen University(深圳大学) Guangming Laboratory(光明实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 针对视频多模态大语言模型中的组合性幻觉问题,提出TriCD框架,通过对比解码和三路径校准机制提升模型在对抗幻觉时的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01125 2026-02-03 cs.CL cs.LG 79%

Long-range Modeling and Processing of Multimodal Event Sequences

多模态事件序列的长程建模与处理

Jichu Li, Yilun Zhong, Zhiting Li, Feng Zhou, Quyu Kong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出了一种基于LLM的多模态时间点过程框架,通过自适应序列压缩解决长上下文问题,提升多模态事件序列的建模与生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21915 2026-02-03 cs.CV 79%

VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models

VideoAesBench: 大规模多模态模型视频审美感知能力评估基准

Yunhao Li, Sijing Wu, Zhilin Gao, Zicheng Zhang, Qi Jia, Huiyu Duan, Xiongkuo Min, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 VideoAesBench通过多样化视频内容和多类型问题评估大规模多模态模型的视频审美感知能力,揭示当前模型在该领域的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00132 2026-02-03 cs.CV 74%

Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time Adaptation

去除伪装,连接领域:通过测试时适应检测转移多模态仇恨视频

Jiao Li, Jian Lang, Xikai Tang, Wenzheng Shu, Ting Zhong, Qiang Gao, Yong Wang, Leiting Chen, Fan Zhou

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 SCANNER通过测试时适应框架,利用仇恨内容中稳定的内核连接源与目标领域,有效应对多模态仇恨视频检测中的语义漂移问题。

Comments Accepted by AAAI2026 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18561 2026-02-03 cs.CV 70%

CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos

CoT-RVS:零样本链式推理视频对象分割

Shiu-hong Kao, Yu-Wing Tai, Chi-Keung Tang

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) National University of Singapore(新加坡国立大学) Dartmouth College(达特茅斯学院)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 CoT-RVS通过零样本链式推理能力,实现了对视频对象的高效分割,无需训练即可处理复杂查询和在线视频流。

Comments Accepted to ICLR 2026. Project page: https://danielshkao.github.io/cot-rvs.html. Code: https://github.com/DanielSHKao/CoT-RVS

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01257 2026-02-03 cs.CV 70%

Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment

通过文本细化与对齐提升点监督时间动作定位

Yunchuan Ma, Laiyun Qing, Guorong Li, Yuqing Liu, Yuankai Qi, Qingming Huang

机构 * University of Chinese Academy of Science, Beijing,100190, China(中国科学院大学,北京) Macquarie University(麦考瑞大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出TRA框架,通过文本细化与对齐提升点监督时间动作定位的精度和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01004 2026-02-03 cs.CV 70%

SRVAU-R1: Enhancing Video Anomaly Understanding via Reflection-Aware Learning

SRVAU-R1: 通过反射感知学习增强视频异常理解

Zihao Zhao, Shengting Cao, Muchao Ye

机构 * The University of Iowa(艾奥瓦大学) Knox College(克诺克斯学院)

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 SRVAU-R1通过引入反射感知学习框架,提升视频异常理解的推理能力和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01683 2026-02-03 cs.CV cs.AI 62%

FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding

FreshMem: 基于大脑的频率-空间混合内存用于流视频理解

Kangcong Li, Peng Ye, Lin Zhang, Chao Wang, Huafeng Qin, Tao Chen

机构 * College of Future Information Technology, Fudan University, Shanghai, China(复旦大学未来信息科技学院) Shanghai Innovation Institute, Shanghai, China(上海创新研究院) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室) The Chinese University of Hong Kong, Hong Kong, China(香港中文大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 FreshMem通过频率-空间混合内存网络,提升流视频理解的连续感知能力,实现短期保真与长期一致性的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04282 2026-02-03 cs.CV cs.LG 57%

Inference-time Stochastic Refinement of GRU-Normalizing Flow for Real-time Video Motion Transfer

推理时GRU归一化流的随机细化用于实时视频运动转移

Tasmiah Haque, Srinjoy Das

机构 * Department of Industrial and Management Systems Engineering(工业与管理系统工程系) School of Mathematical and Data Sciences(数学与数据科学学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种在推理时结合随机采样和GRU-NF的改进方法,以提升实时视频运动转移中未来预测的多样性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06051 2026-02-03 cs.CV 57%

VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement Learning

VQAThinker: 通过强化学习探索通用且可解释的视频质量评估

Linhan Cao, Wei Sun, Weixia Zhang, Xiangyang Zhu, Jun Jia, Kaiwei Zhang, Dandan Zhu, Guangtao Zhai, Xiongkuo Min

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 VQAThinker通过强化学习方法,结合大模型和规则引导算法,提升视频质量评估的泛化能力和可解释性。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00536 2026-02-03 cs.CV 57%

SADER: Structure-Aware Diffusion Framework with DEterministic Resampling for Multi-Temporal Remote Sensing Cloud Removal

SADER:一种结构感知扩散框架,用于多时相遥感云去除

Yifan Zhang, Qian Chen, Yi Liu, Wengen Li, Jihong Guan

机构 * College of Literature, Science, and the Arts, University of Michigan(文学、科学与艺术学院,密歇根大学) School of Computer Science and Technology, Tongji University(计算机科学与技术学院,同济大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 SADER通过结构感知扩散框架,结合时间融合和混合注意力机制,有效解决多时相遥感云去除问题,提升云去除效果和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00440 2026-02-03 cs.CV cs.LG cs.RO 57%

DISK: Dynamic Inference SKipping for World Models

DISK:用于世界模型的动态推理跳过

Anugunj Naman, Gaibo Zhang, Ayushman Singh, Yaguang Zhang

机构 * Purdue University, West Lafayette, United States(普渡大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

AI总结 DISK通过动态推理跳过技术,在无需训练的情况下提升自回归世界模型的视频和轨迹预测效率,同时保持预测精度和视觉质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02396 2026-02-03 cs.RO cs.LG 50%

PRISM: Performer RS-IMLE for Single-pass Multisensory Imitation Learning

PRISM:基于单次传递的RS-IMLE单次传递多感官模仿学习

Amisha Bhaskar, Pratap Tokekar, Stefano Di Cairano, Alexander Schperberg

专题命中 视频多模态 :multimodal(abstract)

AI总结 PRISM通过单次传递的RS-IMLE方法,实现高效的多感官模仿学习,优于扩散策略,提升成功率并减少轨迹 jerk。

Comments 10 pages main text and 4 figures, and 11 pages appendix and 10 figures, total 21 pages and 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02086 2026-02-03 eess.SP cs.HC 50%

Neurophysiological effects of museum modalities on emotional engagement with real artworks

博物馆模态对真实艺术品情感参与的神经生理效应

Chen Feng, Sébastien Lugan, Karine Lasaracina, Midori Sugaya, Benoît Macq

专题命中 视频多模态 :multimodal(abstract)

AI总结 本研究通过EEG分析发现,不同数字解释性内容模态对艺术欣赏的情感参与产生不同影响,为博物馆优化解释媒体提供了新方向。

Comments 7 pages, 4 figures - \c{opyright}IEEE EmotionSense 2026/PerCom 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01156 2026-02-03 cs.LG cs.RO 50%

PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement Learning

PolicyFlow: 在强化学习中使用连续归一化流进行策略优化

Shunpeng Yang, Ben Liu, Hua Chen

机构 * Hong Kong University of Science and Technology(香港科技大学) Southern University of Science and Technology(南方科技大学) Zhejiang University-University of Illinois Urbana-Champaign Institute(浙江大学-伊利诺伊大学厄巴纳-香槟分校联合研究所) LimX Dynamics

专题命中 视频多模态 :multimodal(abstract)

AI总结 PolicyFlow是一种基于连续归一化流的在线强化学习算法,通过减少似然计算开销实现稳定训练,并通过布朗正则化器促进多样化行为。

Comments Submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 9 篇

2506.09114 2026-02-03 cs.LG 82%

TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval

TRACE: 在上下文中对时间序列进行建模以实现多模态嵌入与检索

Jialin Chen, Ziyu Zhao, Gaukhar Nurbek, Aosong Feng, Ali Maatouk, Leandros Tassiulas, Yifeng Gao, Rex Ying

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

AI总结 TRACE通过在上下文中对时间序列进行建模,实现多模态嵌入与检索,提升下游任务的预测精度和可解释性,同时作为强大的独立编码器优化上下文感知表示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02004 2026-02-03 cs.CV cs.AI 81%

ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning

ClueTracer: 问题到视觉线索追踪用于无训练 hallucination 抑制在多模态推理

Gongli Xi, Kun Wang, Zeming Gao, Huahui Yi, Haolang Lu, Ye Tian, Wendong Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学) West China Biomedical Big Data Center(西京生物大数据中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 ClueTracer通过问题到视觉线索追踪,无训练抑制多模态推理中的幻觉,提升推理和非推理任务性能。

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01561 2026-02-03 cs.CV cs.AI 81%

Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd

多模态UNcommonsense:从奇特到普通和从普通到奇特

Yejin Son, Saejin Kim, Dongjun Min, Younjae Yu

机构 * Yonsei University(延世大学) Seoul National University(首尔国立大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 多模态UNcommonsense通过R-ICL框架提升模型在非典型场景下的推理能力,实现从奇特到普通和从普通到奇特的转换。

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01594 2026-02-03 cs.CV 79%

UV-M3TL: A Unified and Versatile Multimodal Multi-Task Learning Framework for Assistive Driving Perception

UV-M3TL: 一种统一且多功能的多模态多任务学习框架用于辅助驾驶感知

Wenzhuo Liu, Qiannan Guo, Zhen Wang, Wenshuo Wang, Lei Yang, Yicheng Qiao, Lening Wang, Zhiwei Li, Chen Lv, Shanghang Zhang, Junqiang Xi, Huaping Liu

机构 * Energy and Transportation Domain, Beijing Institute of Technology(能源与交通领域,北京理工大学) State Key Laboratory of Intelligent Technology and Systems and Department of Computer Science and Technology, Tsinghua University(智能技术与系统国家重点实验室和清华大学计算机科学与技术系) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学) School of Transportation Science and Engineering and the State Key Lab of Intelligent Transportation System, Beihang University(交通运输科学与工程学院和智能交通系统国家重点实验室,北京航空航天大学) Beijing University of Chemical Technology(北京化工大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 UV-M3TL通过双分支结构和自适应损失机制,实现多模态多任务学习,提升辅助驾驶感知的性能与多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21033 2026-02-03 cs.SD cs.AI 70%

SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization

SupCLAP:通过支持向量正则化控制音频-文本对比学习中的优化轨迹漂移

Jiehui Luo, Yuguo Yin, Yuxin Xie, Jinghan Ru, Xianwei Zhuang, Minghua He, Aofan Liu, Zihan Xiong, Dongchao Yang

机构 * Peking University(北京大学) Central Conservatory of Music(中央音乐学院) The Chinese University of Hong Kong(香港中文大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 SupCLAP通过支持向量正则化有效控制音频-文本对比学习中的优化轨迹漂移,提升多模态学习的稳定性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00131 2026-02-03 cs.CV cs.RO 70%

PovNet+: A Deep Learning Architecture for Socially Assistive Robots to Learn and Assist with Multiple Activities of Daily Living

PovNet+: 一种深度学习架构用于社交辅助机器人学习和协助多种日常活动

Fraser Robinson, Souren Pashangpour, Matthew Lisondra, Goldie Nejat

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 PovNet+是一种多模态深度学习架构,用于社交辅助机器人识别多种日常活动并主动发起辅助行为,提升了ADL分类准确率和人机交互能力。

Comments Submitted to Advanced Robotics (Taylor & Francis)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00621 2026-02-03 cs.CV 57%

Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering

通过对比神经元引导实现大型视觉语言模型中幻觉分析与缓解

Guangtao Lyu, Xinyi Cheng, Qi Liu, Chenghao Xu, Jiexi Yan, Muli Yang, Fen Fang, Cheng Deng

机构 * School of Electronic Engineering, Xidian University, Xi'an, China(西安电子科技大学电子工程学院) School of Computer Science and Technology, Xidian University, Xi'an, China(西安电子科技大学计算机科学与技术学院) College of Computer and Information, Hohai University, Nanjing, China(河海大学计算机与信息学院) Institute for Infocomm Research, A*STAR, Singapore(新加坡资讯研究院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 通过对比神经元引导方法,分析并缓解大型视觉语言模型中的幻觉问题,提升视觉表示的稳健性和语义基础性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00547 2026-02-03 cs.LG cs.AI 57%

Contrastive Domain Generalization for Cross-Instrument Molecular Identification in Mass Spectrometry

对比域泛化用于质谱中跨仪器分子识别

Seunghyun Yoo, Sanghong Kim, Namkyung Yoon, Hwangnam Kim

机构 * School of Electrical Engineering, Korea University, Seoul 02841, Korea(韩国大学电气工程学院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种跨模态对齐框架,通过将质谱直接映射到预训练化学语言模型的分子结构嵌入空间,提升跨仪器分子识别的泛化能力。

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05266 2026-02-03 cs.AR cs.CL cs.LG 57%

Understanding and Mitigating Errors of LLM-Generated RTL Code

理解并缓解LLM生成的RTL代码错误

Jiazheng Zhang, Cheng Liu, Long Cheng, Xiaowei Li, Huawei Li

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 本文提出基于LLM的框架,通过检索增强生成、规则检查、多模态转换和迭代仿真调试,显著提升了RTL代码生成的准确性。

Comments Accepted by IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态生成 25 篇

2602.02204 2026-02-03 cs.DC 88%

vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models

vLLM-Omni: 任何到任何多模态模型的完全解耦服务

Peiqi Yin, Jiangyun Zhu, Han Gao, Chenguang Zheng, Yongxiang Huang, Taichang Zhou, Ruirui Yang, Weizhi Liu, Weiqing Chen, Canlin Guo, Didan Deng, Zifeng Mo, Cong Wang, James Cheng, Roger Wang, Hongsheng Liu

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(title,abstract)

AI总结 vLLM-Omni通过解耦服务系统优化多模态模型的高效服务,显著提升作业完成时间效率。

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02140 2026-02-03 cs.CL 83%

Quantifying the Gap between Understanding and Generation within Unified Multimodal Models

量化统一多模态模型中理解与生成之间的差距

Chenlong Wang, Yuhang Chen, Zhihan Hu, Dongping Chen, Wenhu Chen, Sarah Wiegreffe, Tianyi Zhou

机构 * University of Maryland(马里兰大学) University of Waterloo(滑铁卢大学) MBZUAI

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本研究通过GapEval基准量化统一多模态模型中理解与生成能力的差距,揭示当前模型仅实现表面层次的统一,而非深度认知融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25178 2026-02-03 cs.CV cs.AI cs.LG 81%

GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs

GHOST:诱导幻觉的多模态大语言模型图像生成

Aryan Yazdan Parast, Parsa Hosseini, Hesam Asadollahzadeh, Arshia Soltani Moakhar, Basim Azam, Soheil Feizi, Naveed Akhtar

机构 * The University of Melbourne(墨尔本大学) University of Maryland(马里兰大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 GHOST通过优化隐蔽令牌生成诱导幻觉的图像,评估多模态大语言模型的可靠性,并发现高幻觉成功率及可转移漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01901 2026-02-03 cs.CV 79%

Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model

Q Cache:视觉注意力在少于一半的解码层中具有价值用于多模态大语言模型

Jiedong Zhuang, Lu Lu, Ming Dai, Rui Hu, Jian Chen, Qiang Liu, Haoji Hu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 Q Cache通过跨层共享相似注意力模式,减少多模态大语言模型中的KV缓存使用量,提升吞吐量并保持性能

Comments Accepted by AAAI26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00960 2026-02-03 cs.LG cs.AI cs.CE stat.CO stat.ML 79%

Multimodal Scientific Learning Beyond Diffusions and Flows

多模态科学学习超越扩散与流

Leonardo Ferreira Guilhoto, Akshat Kaushal, Paris Perdikaris

机构 * Graduate Group in Applied Mathematics and Computational Science(应用数学与计算科学联合研究生组) Department of Computer and Information Science(计算机与信息科学系) Department of Mechanical Engineering and Applied Mechanics(机械工程与应用力学系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出混合密度网络作为多模态科学学习中更高效、更稳定的替代方法,实现对科学问题中多模态不确定性的有效建模与解决。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00849 2026-02-03 cs.LG cs.AI cs.NA math.NA 79%

RMFlow: Refined Mean Flow by a Noise-Injection Step for Multimodal Generation

RMFlow:通过噪声注入步骤细化均流以实现多模态生成

Yuhao Huang, Shih-Hsin Wang, Andrea L. Bertozzi, Bao Wang

机构 * Department of Mathematics and Scientific Computing and Imaging (SCI) Institute University of Utah(数学与科学计算及成像学院(SCI)院,犹他大学) Department of Mathematics, UCLA(数学系,加州大学洛杉矶分校)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 RMFlow通过引入噪声注入步骤,改进均流模型,实现高效多模态生成,仅需单次功能评估即可达到接近最先进的性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏