arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46618 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4790 篇

2511.00405 2026-03-03 cs.LG cs.AI 79%

UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings

UME-R1: 探索基于推理的生成多模态嵌入

Zhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou, Jinsong Su

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) WeChat AI, Tencent Inc, China(腾讯公司微信AI部门) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 UME-R1通过生成嵌入和强化学习提升多模态嵌入性能,实现判别与生成嵌入的互补,展现生成嵌入在推理和下游任务中的优势。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03905 2026-03-03 cs.CV 79%

EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation

EchoMimicV3:13亿参数足矣实现统一的多模态和多任务人类动画

Rang Meng, Yan Wang, Weipeng Wu, Ruobing Zheng, Yuming Li, Chenguang Ma

机构 * Terminal Technology Department, Alipay, Ant Group(蚂蚁集团支付宝终端技术部)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 EchoMimicV3通过统一多任务和多模态人类动画,以13亿参数实现高效且稳定的多任务和多模态动画生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23937 2026-03-02 cs.RO cs.CV 79%

Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos

通过真实世界室内游览视频的多模态事件知识增强视觉语言导航

Haoxuan Xu, Tianfu Li, Wenbo Chen, Yi Liu, Xingxing Zuo, Yaoxian Song, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学) Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(马尔代夫 bin Zayed 大学人工智能学院) Hangzhou City University(杭州城市学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于多模态事件知识的视觉语言导航增强方法,通过构建大规模时空知识图谱并结合层次检索机制,提升长视界推理和粗粒度指令处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08578 2026-03-02 cs.CV 79%

Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities

多模态知识蒸馏用于抗缺失模态的自体视觉动作识别

Maria Santos-Villafranca, Dustin Carrión-Ojeda, Alejandro Perez-Yus, Jesus Bermudez-Cameo, Jose J. Guerrero, Simone Schaub-Meyer

机构 * I3A – University of Zaragoza(萨拉戈萨大学I3A研究中心) Technical University of Darmstadt, Department of Computer Science(达姆施塔特技术大学计算机科学系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 KARMMA通过多模态知识蒸馏实现抗缺失模态的自体视觉动作识别,以轻量模型提升机器人部署效率。

Comments Project Page: https://visinf.github.io/KARMMA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22920 2026-02-27 cs.CV 79%

OSDaR-AR: Enhancing Railway Perception Datasets via Multi-modal Augmented Reality

OSDaR-AR: 通过多模态增强现实提升铁路感知数据集

Federico Nesti, Gianluca D'Amico, Mauro Marinoni, Giorgio Buttazzo

机构 * Department of Excellence in Robotics & AI, Scuola Superiore Sant’Anna(机器人与人工智能卓越部门,圣安娜高等学院) Simulatrix MV srl(Simulatrix MV公司)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出OSDaR-AR数据集,通过多模态增强现实技术整合逼真虚拟对象,提升铁路感知任务的数据质量与真实性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11910 2026-02-26 cs.CV 79%

Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models

看清森林与树木:面向长视频多模态语言模型的查询感知分词器

Siyou Li, Huanan Wu, Juexi Shao, Yinghao Ma, Yujian Gan, Yihao Luo, Yuwei Wang, Dong Nie, Lu Wang, Wenqing Wu, Le Zhang, Massimo Poesio, Juntao Yu

机构 * Queen Mary University of London(伦敦女王学院) University of Sheffield(谢菲尔德大学) Imperial College London(伦敦帝国学院) Pengcheng Laboratory(鹏城实验室) Meta Inc(Meta公司) Meituan Inc(美团公司) Nanjing University of Science(南京理工大学) University of Birmingham(伯明翰大学) Utrecht University(乌得勒支大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 QTSplus是一种轻量高效的视觉token选择模块,通过动态选择重要视觉证据提升长视频多模态语言模型的性能,显著降低计算成本并提高处理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19063 2026-02-24 cs.CV 79%

Direction-aware 3D Large Multimodal Models

具有方向感知的3D大多模态模型

Quan Liu, Weihao Xuan, Junjue Wang, Naoto Yokoya, Ling Shao, Shijian Lu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出方向感知的3D大多模态模型,通过补充自身姿态并改进点云数据对齐,提升3D多模态模型的性能和通用性。

Comments In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17252 2026-02-20 cs.CV cs.SY eess.IV eess.SY 79%

A Multi-modal Detection System for Infrastructure-based Freight Signal Priority

基于基础设施的货运信号优先的多模态检测系统

Ziyan Zhang, Chuheng Wei, Xuanpeng Zhao, Siyan Li, Will Snyder, Mike Stas, Peng Hao, Kanok Boriboonsomsin, Guoyuan Wu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于激光雷达和摄像头的多模态货运车辆检测系统,通过混合传感架构和卡尔曼滤波实现稳定实时性能,用于支持基于基础设施的货运信号优先应用。

Comments 12 pages, 15 figures. Accepted at ICTD 2026. Final version to appear in ASCE Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10902 2026-02-18 cs.CL 79%

Multimodal Peer Review Simulation with Actionable To-Do Recommendations for Community-Aware Manuscript Revisions

多模态同行评审模拟:面向社区意识的论文修订可操作待办推荐系统

Mengze Hong, Di Jiang, Weiwei Zhao, Yawen Li, Yihang Wang, Xinyuan Luo, Yanjie Sun, Chen Jason Zhang

机构 * Hong Kong Polytechnic University(香港理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Independent Researcher(独立研究者)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本研究提出一个多模态同行评审模拟系统,通过整合文本和视觉信息,利用检索增强生成技术提升评审质量,并生成可操作的待办列表,以提高论文修订的效率和准确性。

Comments Accepted by TheWebConf 2026 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21842 2026-02-17 cs.CV cs.CR 79%

Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?

模态失语:统一多模态模型能否从记忆中描述图像?

Michael Aerni, Joshua Swanson, Kristina Nikolić, Florian Tramèr

机构 * Michael Aerni(独立研究者) Joshua Swanson(独立研究者) Kristina Nikolić(独立研究者) Florian Tramèr(独立研究者)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 研究发现统一多模态模型在视觉记忆与文本表达间存在系统性缺陷,导致安全框架可能因单一模态防护而失效。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12593 2026-02-16 cs.IR cs.AI 79%

RQ-GMM: Residual Quantized Gaussian Mixture Model for Multimodal Semantic Discretization in CTR Prediction

RQ-GMM:用于CTR预测的多模态语义离散化残差量化高斯混合模型

Ziye Tong, Jiahao Liu, Weimin Zhang, Hongji Ruan, Derick Tang, Zhanpeng Zeng, Qinsong Zeng, Peng Zhang, Tun Lu, Ning Gu

机构 * Tencent(腾讯) Fudan University(复旦大学) Beijing Jiaotong University(北京交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 RQ-GMM通过残差量化高斯混合模型提升CTR预测中多模态语义离散化效果,实现代码本利用和重建准确性的显著提升。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09638 2026-02-11 cs.CV 79%

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

VideoAfford: 通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding

Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA, Yiming Zhong, Wenti Yin, Yuhao Liu, Zhiqing Cui, Jiahao Yuan, Lu Dai, Zhiyuan Ma, Hui Xiong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 VideoAfford通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding,结合动态交互先验和空间感知损失函数,提升机器人操作的可操作区域识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08861 2026-02-10 cs.CV 79%

TiFRe: Text-guided Video Frame Reduction for Efficient Video Multi-modal Large Language Models

TiFRe: 基于文本的视频帧减少用于高效视频多模态大语言模型

Xiangtian Zheng, Zishuo Wang, Yuxin Peng

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 TiFRe通过文本引导的帧采样和帧匹配机制,有效减少视频输入帧数,同时保留关键信息,提升视频多模态大语言模型的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13564 2026-02-10 cs.CV 79%

State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models

具有门控注意力和可学习采样的状态空间分层压缩用于大多模态模型中的小时级视频理解

Geewook Kim, Minjoon Seo

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于状态空间模型和门控注意力的高效压缩方法,用于减少大模型中小时级视频的token消耗,同时保持性能。

Comments AAAI 2026 (Oral). Project page: https://github.com/naver-ai/mambamia

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13928 2026-02-05 cs.CV cs.IR 79%

LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts

LoVR:一种多模态背景下长视频检索的基准

Qifeng Cai, Hao Liang, Zhaoyang Han, Hejun Dong, Meiyi Qiang, Ruichuan An, Quanqing Xu, Bin Cui, Wentao Zhang

机构 * East China Normal University Shanghai China Peking University \& Zhongguancun Academy Beijing China Huazhong University of Science Beihang University Beijing China Peking University Beijing China East China Normal University Peking University \& Zhongguancun Academy Beihang University Peking University

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 LoVR是一个针对长视频检索的多模态基准,通过高质量标注和细粒度数据提升视频理解挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02453 2026-02-04 cs.AI 79%

Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling

用漫画思考:通过结构化视觉叙事增强多模态推理

Andong Chen, Wenxin Zhu, Qiuyu Ding, Yuchen Song, Muyun Yang, Tiejun Zhao

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出用漫画作为中间视觉表示,通过结构化视觉叙事提升多模态推理效率和性能。

Comments Working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01125 2026-02-03 cs.CL cs.LG 79%

Long-range Modeling and Processing of Multimodal Event Sequences

多模态事件序列的长程建模与处理

Jichu Li, Yilun Zhong, Zhiting Li, Feng Zhou, Quyu Kong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出了一种基于LLM的多模态时间点过程框架,通过自适应序列压缩解决长上下文问题,提升多模态事件序列的建模与生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21915 2026-02-03 cs.CV 79%

VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models

VideoAesBench: 大规模多模态模型视频审美感知能力评估基准

Yunhao Li, Sijing Wu, Zhilin Gao, Zicheng Zhang, Qi Jia, Huiyu Duan, Xiongkuo Min, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 VideoAesBench通过多样化视频内容和多类型问题评估大规模多模态模型的视频审美感知能力,揭示当前模型在该领域的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22039 2026-01-30 cs.CV 79%

Understanding Multimodal Complementarity for Single-Frame Action Anticipation

理解单帧动作预测中的多模态互补性

Manuel Benavent-Lledo, Konstantinos Bacharidis, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez

机构 * Department of Computer Technology, University of Alicante(阿尔瓦雷斯大学计算机技术系) Institute of Computer Science, FORTH(福蒂研究所) Computer Science Department, University of Crete(克里特大学计算机科学系) Department of Management, Science and Technology, Hellenic Mediterranean University(希腊地中海大学管理、科学与技术系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究通过单帧动作预测框架AAG+,探索多模态互补性对动作预测的影响,验证单帧信息在动作预测中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17123 2026-01-27 cs.HC cs.CV cs.RO 79%

Acoustic Field Video for Multimodal Scene Understanding

用于多模态场景理解的声学场视频

Daehwa Kim, Chris Harrison

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出声学场视频作为多模态场景理解的新输入方式,通过整合空间声学数据显著提升视觉-语言模型的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19529 2026-01-21 cs.CV 79%

Vidi2.5: Large Multimodal Models for Video Understanding and Creation

Vidi2.5:大型多模态模型用于视频理解和创作

Vidi Team, Chia-Wen Kuo, Chuang Huang, Dawei Du, Fan Chen, Fanding Lei, Feng Gao, Guang Chen, Haoji Zhang, Haojun Zhao, Jin Liu, Jingjing Zhuge, Lili Fang, Lingxi Zhang, Longyin Wen, Lu Guo, Lu Xu, Lusha Li, Qihang Fan, Rachel Deng, Shaobo Fang, Shu Zhang, Sijie Zhu, Stuart Siew, Weiyan Tao, Wen Zhong, Xiaohui Shen, Xin Gu, Ye Yuan, Yicheng He, Yiming Cui, Zhenfang Chen, Zhihua Wu, Zuhua Lin

机构 * ByteDance Inc.(字节跳动公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 Vidi2.5通过细粒度时空定位和复杂情节推理模型,提升了视频理解和创作的多模态能力,同时在多个基准测试中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12330 2026-01-21 cs.LG cs.AI 79%

IceWatch: Forecasting Glacial Lake Outburst Floods (GLOFs) using Multimodal Deep Learning

IceWatch: 使用多模态深度学习预测冰湖溃决洪水(GLOFs)

Zuha Fatima, Muhammad Anser Sohaib, Muhammad Talha, Ayesha Kanwal, Sidra Sultana, Nazia Perwaiz

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 IceWatch通过多模态深度学习结合空间和时间视角,实现对冰湖溃决洪水的高效预测,提升预警系统的可靠性与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11990 2026-01-21 cs.CV 79%

DAOS: A Multimodal In-cabin Behavior Monitoring with Driver Action-Object Synergy Dataset

DAOS: 一种基于驾驶员行为与物体协同的多模态舱内行为监测数据集

Yiming Li, Chen Cai, Tianyi Liu, Dan Lin, Wenqian Wang, Wenfei Liang, Bingbing Li, Kim-Hui Yap

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院) College of Computer Science and Technology, Harbin Engineering University(哈尔滨工程大学计算机科学与技术学院) Singapore University of Technology and Design (SUTD), Singapore(新加坡科技设计大学) Continental Automotive Singapore Pte. Ltd.

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 DAOS数据集通过多模态数据和动作-物体关系网络提升驾驶员行为识别的准确性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02064 2026-01-16 cs.CV 79%

RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video

RTV-Bench: 通过实时视频对多模态大语言模型的连续感知、理解和推理进行基准测试

Shuhang Xun, Sicheng Tao, Jungang Li, Yibo Shi, Zhixin Lin, Zhanhui Zhu, Yibo Yan, Hanqian Li, Linghao Zhang, Shikang Wang, Yixin Liu, Hanbo Zhang, Ying Ma, Xuming Hu

机构 * HIT(哈尔滨工业大学) HKUST (GZ)(香港科技大学(广州)) HKUST(香港科技大学) XJTU(西安交通大学) SDU(山东大学) CityU(城市大学) HUST(华中科技大学)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 RTV-Bench通过实时视频对多模态大语言模型的连续感知、理解和推理能力进行细粒度基准测试,揭示了实时模型在长时段视频处理中的性能优势与局限。

Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track;

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08408 2026-01-14 cs.CV cs.RO 79%

Edge-Optimized Multimodal Learning for UAV Video Understanding via BLIP-2

边缘优化的多模态学习用于无人机视频理解 via BLIP-2

Yizhan Feng, Hichem Snoussi, Jing Teng, Jian Liu, Yuyang Wang, Abel Cherouat, Tian Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于BLIP-2的轻量级多模态学习平台,通过集成YOLO模型实现无人机视频理解的高效处理与多任务适应。

Comments The Tenth International Conference on Data Mining and Big Data (DMBD'2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08240 2026-01-14 eess.IV cs.CV 79%

Temporal-Enhanced Interpretable Multi-Modal Prognosis and Risk Stratification Framework for Diabetic Retinopathy (TIMM-ProRS)

时间增强的可解释多模态预后和风险分层框架用于糖尿病视网膜病变(TIMM-ProRS)

Susmita Kar, A S M Ahsanul Sarkar Akib, Abdul Hasib, Samin Yaser, Anas Bin Azim

机构 * Department of Robotics, Robo Tech Valley, Dhaka, Bangladesh(机器人系,罗布科技谷,达卡,孟加拉国)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 TIMM-ProRS通过融合视网膜图像和时间生物标志物,实现糖尿病视网膜病变的多模态和时间动态分析,达到97.8%的准确率,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22393 2026-01-13 cs.CV cs.LG 79%

Gems: Group Emotion Profiling Through Multimodal Situational Understanding

GEMS: 通过多模态情境理解进行群体情绪分析

Anubhav Kataria, Surbhi Madan, Shreya Ghosh, Tom Gedeon, Abhinav Dhall

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 GEMS通过多模态情境理解实现群体情绪分析,预测个体、群体和事件层面的情绪,提供更细粒度和整体的分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04709 2026-01-09 cs.AI 79%

Bridging Temporal and Textual Modalities: A Multimodal Framework for Automated Cloud Failure Root Cause Analysis

弥合时序与文本模态:一种多模态框架用于自动化云故障根本原因分析

Gijun Park

机构 * Okestro

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种多模态框架,通过融合时间序列与文本数据,提升云故障根本原因分析的自动化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00123 2026-01-09 cs.CV 79%

A Spatially Masked Adaptive Gated Network for multimodal post-flood water extent mapping using SAR and incomplete multispectral data

一种空间掩码自适应门控网络用于利用SAR和不完整多光谱数据的多模态洪水后水 extent 映射

Hyunho Lee, Wenwen Li

机构 * School of Geographical Sciences and Urban Planning, Arizona State University(地理科学与城市规划学院,亚利桑那州立大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出SMAGNet,一种多模态深度学习模型,通过融合SAR和不完整多光谱数据,提升洪水后水 extent 映射的准确性和鲁棒性。

Comments 50 pages, 12 figures, 6 tables

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing, 232, 492-508, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02968 2026-01-07 cs.AI 79%

Rationale-Grounded In-Context Learning for Time Series Reasoning with Multimodal Large Language Models

基于理由的上下文学习用于多模态大语言模型的时间序列推理

Qingxiang Liu, Zhiqing Cui, Xiaoliang Luo, Yuqian Wu, Zhuoyang Jiang, Huaiyu Wan, Sheng Sun, Lvchun Wang, Wei Yu, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) China Mobile (Jiangxi) Virtual Reality Technology Co., Ltd.(中国移动(江西)虚拟现实技术有限公司) School of Computer and Information Technology, Beijing Jiaotong University(北京交通大学计算机与信息学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本研究提出RationaleTS方法,通过基于理由的上下文学习提升多模态大语言模型在时间序列推理中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏