arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4772 篇

2208.08090 2022-08-18 cs.CV cs.MM 84%

Progressive Cross-modal Knowledge Distillation for Human Action Recognition

Jianyuan Ni, Anne H. H. Ngu, Yan Yan

专题命中 视频多模态 :cross-modal(title);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06085 2022-04-12 cs.CV cs.CL 84%

On Pursuit of Designing Multi-modal Transformer for Video Grounding

Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, Yuexian Zou

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by Conference on Empirical Methods in Natural Language Processing (EMNLP 2021, Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02566 2022-04-07 cs.CL cs.MM 84%

Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension

Huibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang, Yufan Li, Ning jiang, Xin wei, Zhenglu Yang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.MM

Comments Accepted by ACL-2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.07175 2021-05-18 cs.CV cs.MM 84%

Cross-Modal Progressive Comprehension for Referring Segmentation

Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei, Bo Li, Guanbin Li

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

Comments Accepted by TPAMI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.03848 2021-03-04 cs.CV cs.CL 84%

Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers

Shijie Geng, Peng Gao, Moitreya Chatterjee, Chiori Hori, Jonathan Le Roux, Yongfeng Zhang, Hongsheng Li, Anoop Cherian

专题命中 视频多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL

Comments Accepted at AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.03049 2020-02-18 cs.CV cs.CL 84%

Video Question Generation via Cross-Modal Self-Attention Networks Learning

Yu-Siang Wang, Hung-Ting Su, Chen-Hsi Chang, Zhe-Yu Liu, Winston H. Hsu

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICASSP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.04268 2019-05-23 cs.IR cs.CV cs.MM 84%

Learning Shared Semantic Space with Correlation Alignment for Cross-modal Event Retrieval

Zhenguo Yang, Zehang Lin, Peipei Kang, Jianming Lv, Qing Li, Wenyin Liu

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

Comments 22 pages, submitted to ACM Transactions on Multimedia Computing Communications and Applications(ACM TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.07023 2018-12-19 cs.CL cs.CV 84%

From FiLM to Video: Multi-turn Question Answering with Multi-modal Context

Dat Tien Nguyen, Shikhar Sharma, Hannes Schulz, Layla El Asri

专题命中 视频多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL

Comments Accepted for an Oral presentation at the DSTC7 workshop at AAAI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1512.00818 2015-12-17 cs.CV cs.CL cs.LG 84%

Zero-Shot Event Detection by Multimodal Distributional Semantic Embedding of Videos

Mohamed Elhoseiny, Jingen Liu, Hui Cheng, Harpreet Sawhney, Ahmed Elgammal

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments To appear in AAAI 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1308.1150 2013-08-07 cs.MM cs.CV 84%

Multimodal Approach for Video Surveillance Indexing and Retrieval

Ali Wali, Adel M. Alimi

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.MM

Comments 7 pages

Journal ref Journal of Intelligent Computing, Volume: 1, Issue: 4 (December 2010), Page: 165-175

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10783 2023-09-20 cs.CV cs.AI cs.CL 83%

Language as the Medium: Multimodal Video Classification through text only

Laura Hanu, Anita L. Verő, James Thewlis

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)

Comments Accepted at "What is Next in Multimodal Foundation Models?" (MMFM) workshop at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26067 2026-08-27 cs.CV 新提交 83%

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

StreamPI:面向视觉-语言-动作模型的流式多模态时序建模

Zhe Liu, Jinghua Hou, Yuxiang Lu, Zhenya Yang, Xianzhe Fan, Junwei Luo, Junyi Li, Ruihua Han, Zhi Hou, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) ACE Robotics(ACE机器人公司)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出StreamPI框架,为单帧VLA模型赋予时序推理能力,通过指令锚定时序建模、随机间隔流式训练等方法,在真实机器人与仿真基准任务上性能优于pi0.5。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14854 2026-08-18 cs.CV 新提交 83%

Zero-MELO: Test-Time Evidence Calibration with Multimodal LLMs for Zero-Shot Micro-Gesture Recognition

Zero-MELO:基于多模态大语言模型的测试时证据校准用于零样本微手势识别

Chengyan Wang, Hanliang Xie, Yueyi Yang, Haoyu Chen

机构 * University of Oulu(奥卢大学) Peking University(北京大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究针对多模态大语言模型在微手势识别中局部证据不足、分数偏差的瓶颈,提出Zero-MELO框架,结合树搜索、测试时校准与多线索融合,在iMiGUE和MA-52数据集上显著优于Qwen2.5-VL基线。

Comments Accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01045 2026-08-11 cs.AI 版本更新 83%

Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets

Med-CRAFT:通过知识图谱遍历自动构建可解释的多跳视频工作负载

Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li

机构 * Beijing Institute of Technology(北京理工大学) The Hong Kong Polytechnic University(香港理工大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract_cn);分类 cs.AI

AI总结 Med-CRAFT通过知识图谱遍历自动构建可解释的多跳视频工作负载,生成具有细粒度时间选择性和多跳逻辑复杂性的医疗视频推理基准。

Comments 23 pages, 4 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02324 2026-08-04 cs.CV 新提交 83%

Implicit Neural Representations for Multimodal Longitudinal Image Imputation and Interpolation

用于多模态纵向图像插补与插值的隐式神经表示

Sina Wendrich, Lukas Förner, Zoe Reinke, Kartikay Tehlan, Ansgar Berlis, Michael Frühwald, Matthias Wagner, Thomas Wendler

机构 * University Hospital Augsburg(奥格斯堡大学医院) University of Augsburg(奥格斯堡大学) Technical University of Munich(慕尼黑工业大学) Bavarian Cancer Research Center (BZKF)(巴伐利亚癌症研究中心) Swabian Children’s Cancer Center(施瓦本儿童癌症中心)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 该研究针对临床纵向MRI数据缺失等问题,提出条件隐式神经表示模型,经儿科脑肿瘤数据验证,可显著改进插值效果,置信度估计可靠。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03168 2026-07-28 cs.CV cs.LG 版本更新 83%

Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs

Farm-LightSeek:一种以边缘为中心的多模态农业物联网数据分析框架,集成轻量级语言模型

Dawen Jiang, Zhishu Shen, Qiushi Zheng, Tiehua Zhang, Wei Xiang, Jiong Jin

机构 * School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院) School of Engineering, Swinburne University of Technology(斯威本科技大学工程学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) School of Computing, Engineering and Mathematical Sciences, La Trobe University(拉筹伯大学计算、工程与数学科学学院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 针对智能农业面临的挑战,提出Farm-LightSeek框架,将大语言模型与边缘计算集成。通过传感器收集多源数据,在边缘节点进行跨模态推理等。创新包括闭环架构等,实验表明该框架在关键任务中性能可靠,推动了智能实时农业及两者深度集成。

Comments Accepted by IEEE Internet of Things Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20981 2026-07-24 cs.AI 新提交 83%

Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence

超越独立优化:多模态边缘智能中的压缩、混合专家路由和量化交互

Jay Gor, Karm Dave, Akshita Abrol, Rajesh Gupta, Sudeep Tanwar, Zhengkui Wang

机构 * Nirma University(尼玛大学) Singapore Institute of Technology(新加坡理工学院) Marwadi University(马尔瓦迪大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 研究多模态边缘智能中高效推理受多种因素限制,回顾相关模型进展,指出技术间相互影响不能独立优化,介绍关键设计权衡,引入视频MoE模型诊断方法,强调多方面开放研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17712 2026-07-21 cs.AI 新提交 83%

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution

学习检测跨模态否定:潜在表示分析与基于注意力的解决方案

Ali AbuSaleh, Leon Hammerla, Alexander Mehler

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 研究跨模态否定检测难题,提出新型跨模态注意力架构,分析发现文本与视觉否定的不对称性,结合自监督视频表示推进时间否定建模,为多模态系统学习语义对齐表示提供新方法。

Comments This manuscript is an accepted version of the article (published at ICNLP2026). Published in IEEE Xplore, DOI:10.1109/ICNLP69856.2026.11527861 document: https://ieeexplore.ieee.org/abstract/document/11527861

Journal ref 2026 8th International Conference on Natural Language Processing (ICNLP), Xi'an, China, 2026, pp. 613-622

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05971 2026-07-08 cs.MM cs.SD 新提交 83%

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking

通过语义检索和时间重排实现多模态视频到音乐推荐

Seungheon Doh, Minhee Lee, Sangmoon Lee, Ben Sangbae Chon, Juhan Nam

机构 * Graduate School of Culture Technology, KAIST(韩国釜山科学技术大学文化科技研究生院) Gaudio Lab, Inc(Gaudio实验室)

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.MM

AI总结 提出用于视频到音乐推荐的两阶段框架VTMR,第一阶段通过全局嵌入检索语义兼容候选,第二阶段关注时间序列重排。实验显示该框架能提升推荐效果,人类偏好研究表明其在总体偏好上与商业基线相当,音乐质量优于生成基线。

Comments Accepted for publication at The Machine Learning for Audio workshop at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11730 2026-07-07 cs.CV cs.HC cs.LG 版本更新 83%

Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

视频中矛盾/犹豫识别用于个性化数字健康干预

Manuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi, Muhammad Haseeb Aslam, Lorenzo Sia, Nicolas Richet, Marco Pedersoli, Alessandro Lameiras Koerich, Simon L Bacon, Eric Granger

机构 * LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(ETS蒙特利尔大学系统工程系LIVIA实验室) LIVIA, Dept. of Software and IT Engineering, ETS Montreal, Canada(ETS蒙特利尔大学软件与信息工程系LIVIA实验室) Dept. of Health, Kinesiology, & Applied Physiology, Concordia University, Montreal, Canada(康科迪亚大学健康、运动科学与应用生理学系) Montreal Behavioural Medicine Centre, CIUSSS Nord-de-l’Ile-de-Montréal, Canada(蒙特利尔行为医学中心,蒙特利尔北岛卫生与社会服务局)

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文研究了通过深度学习模型在视频中进行矛盾/犹豫识别,以提升数字健康干预的个性化和成本效益,实验基于新的BAH视频数据集,发现需改进多模态模型以准确识别矛盾/犹豫。

Comments 11 pages, 4 figures, ACII 2026. arXiv admin note: substantial text overlap with arXiv:2505.19328

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01667 2026-07-03 cs.CV 新提交 83%

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning

增强视听视频字幕的时间与跨模态对齐

Chen Zhao, Jiajun Ma, Qilong Huang, Tiehan Fan, Hongyu Li, Zhuoliang Kang, Xiaoming Wei, Jian Yang, Ying Tai

机构 * Nanjing University, State Key Laboratory for Novel Software Technology(南京大学,计算机软件新技术国家重点实验室) Meituan(美团)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 提出TCA-Captioner框架,通过观察-检查-纠正迭代策略和密集交互数据集,解决视听字幕中的模态分离与时间不一致问题,实现精准的视听绑定与因果动态建模。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08945 2026-06-23 cs.CV 版本更新 83%

MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment

PIDNet: 逐步隐式解耦网络用于多模态动作质量评估

Qiqi Li, Pengfei Wang, Hongyu Chen, Nenggan Zheng

机构 * Qiushi Academy for Advanced Studies (QAAS), Zhejiang University(浙江大学启斯特先进研究院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Software Technology, Zhejiang University(浙江大学软件学院) State Key Lab of Brain-Machine Intelligence(脑机智能国家重点实验室) Collaborative Innovation Center for Artificial Intelligence by MOE and Zhejiang Provincial Government (ZJU)(教育部-浙江省人工智能协同创新中心) Zhejiang Lab(浙江实验室)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出PIDNet,通过逐步整合多模态信息与全局质量语义,提升多模态动作质量评估的准确性。采用iMambaWave模块和三阶段融合网络,有效解耦模态特定信息并增强特征表示。

Comments 14 pages, 6 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04098 2026-06-23 cs.CV 版本更新 83%

A Physics-Informed, Behavior-Aware Digital Twin for Robust Multimodal Forecasting of Core Body Temperature in Precision Livestock Farming

一种物理知情、行为感知的数字孪生方法,用于精准畜牧业中核心体温的鲁棒多模态预测

Riasad Alvi, Mohaimenul Azam Khan Raiaan, Sadia Sultana Chowa, Arefin Ittesafun Abian, Reem E Mohamed, Md Rafiqul Islam, Yakub Sebastian, Sheikh Izzal Azid, Sami Azam

机构 * Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory(应用人工智能与智能系统实验室) Department of Computer Science and Engineering(计算机科学与工程系) Department of Data Science and Artificial Intelligence(数据科学与人工智能系) Department of Software Systems & Cybersecurity(软件系统与网络安全系) Energy and Resources Institute, Faculty of Science and Technology(能源与资源研究所,科学与技术学院) Faculty of Science and Technology(科学与技术学院) School of Engineering and Energy(工程与能源学院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出物理知情数字孪生框架,结合不确定性感知的专家加权堆叠集成,利用ODE热调节模型、高斯过程、卡尔曼滤波和行为马尔可夫链,融合多模态数据实现奶牛核心体温的2小时提前预测,R²达0.783。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06285 2026-06-05 cs.AI 83%

TRACE: A Temporal Conditional Estimation for Multimodal Time Series Foundation Models

TRACE: 面向多模态时间序列基础模型的时间条件估计

Ziwen Kan, Yishuo Chen, Kecheng Li, Andrew Wen, Xiaomeng Wang, Liwei Wang, Jihao Duan, Song Wang, Hongfang Liu, Tianlong Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 提出TRACE条件估计范式,通过利用可用辅助模态推断缺失目标模态,解决多模态时间序列中的时间错位和部分模态缺失问题,在医疗和情感分析基准上优于现有融合方法。

Comments 5 figures and 5 tables in the main paper, plus appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05997 2026-06-05 cs.CV 83%

Multimodal Sexism Identification and Characterization using Large Language Models and Gradient Boosting

使用大语言模型和梯度提升的多模态性别歧视识别与表征

Kyriakos Chaviaras, Maria Lymperaiou, Athanasios Voulodimos

机构 * Artificial Intelligence and Learning Systems Laboratory(人工智能与学习系统实验室) School of Electrical and Computer Engineering(电气与计算机工程学院) National Technical University of Athens(雅典国家技术大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出基于特征工程和梯度提升回归模型的后融合管道,结合视觉、文本、人口统计、生物特征及LLM语义指标,用于识别和表征模因和短视频中的多模态性别歧视。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24302 2026-05-28 cs.CV 83%

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies

基于Mamba的第一人称视频跨模态动作识别:通过CLS令牌融合策略整合RGB和手部骨架流

Juan Ignacio Bustos Gorostegui, Maria Elena Buemi

机构 * Univ. of Buenos Aires. Faculty of Exact and Natural Sciences. Dept. of Computer Science (DC)(布宜诺斯艾利斯大学。精确与自然科学学院。计算机科学系) CONICET-Univ. of Buenos Aires. Institute of Computer Sciences (ICC)(布宜诺斯艾利斯大学CONICET联合体。计算机科学研究所)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 提出一种基于Mamba的跨模态架构,通过四种CLS令牌融合策略(朴素、平均、加权和基于上下文)整合RGB视频和手部骨架数据,在H2O数据集上平均策略达到最佳性能,Top-1准确率在Tiny配置下提升超10%。

Comments 4 pages , 2 figures , Egovis2026 , CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25409 2026-05-26 cs.CV 83%

MTLLFM: Multimodal-Temporal Laughter Localization: UR-FUNNY-Temporal and SMILE-Temporal Benchmarks with an Adaptive Multimodal Fusion Model

MTLLFM: 多模态时间笑声定位——UR-FUNNY-Temporal和SMILE-Temporal基准数据集与自适应多模态融合模型

Eyal Hanania, Nadav Kirsch, Daniel Arkushin, Jonathan Benvenisti, Amos Bercovich, Elie Zemmour, Sahar Froim

机构 * WSC-Sports(WSC-体育)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

AI总结 针对现有方法无法精确捕捉短暂笑声事件时间边界的问题,本文提出两个完全标注的时间笑声数据集(UR-FUNNY-Temporal和SMILE-Temporal)和一个轻量级弱监督框架,通过固定HuBERT和MAE编码器结合时间softmax池化与自适应模态门控,实现从片段级标签到帧级时间定位,在体育广播数据上达到99% F1和68.1%定位精度,并将下游笑声推理CIDEr提升227%。

Comments Accepted to the Workshop on Affective & Behavior Analysis in-the-wild, CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19950 2026-05-20 cs.CV 83%

AffectVerse: Emotional World Models for Multimodal Affective Computing

AffectVerse: 多模态情感计算中的情感世界模型

Bo Zhao, Fanghua Ye, Yixin Ji, Sicheng Zhao, Xiaojiang Peng, Zitong YU

机构 * Great Bay University(大湾大学) Tencent(腾讯) Tsinghua University(清华大学) Shenzhen Technology University(深圳技术大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本研究提出AffectVerse,一种基于Qwen2.5-Omni的多模态情感计算模型,通过引入情感世界模块实现短期潜在情感预测,利用未来预测作为自监督信号,提高了情感计算的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13026 2026-05-15 cs.CV 83%

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

REVISOR:超越文本反思,迈向长视频理解的多模态反思推理

Jiaze Li, Hao Yin, Wenhui Tan, Jingyang Chen, Boshen Xu, Yuxun Qu, Yijing Chen, Jianzhong Ju, Zhenbo Luo, Jian Luan

机构 * MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus实验室) Renmin University of China(中国人民大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 REVISOR通过多模态反思框架提升长视频理解能力,解决纯文本反思在动态视觉输入和跨模态交互上的不足,采用Dual Attribution Decoupled Reward机制增强因果对齐,无需额外监督微调即在多个基准测试中取得优异成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13248 2026-05-14 eess.SP cs.AI 83%

Compact Latent Manifold Translation: A Parameter-Efficient Foundation Model for Cross-Modal and Cross-Frequency Physiological Signal Synthesis

紧凑的潜在流形翻译:一种参数高效的基础模型用于跨模态和跨频率生理信号合成

Bo Cui, Xiaowen Song, Yaowen Zhang, Shunzhe Zhang, B. J. F. van Beijnum, Monique Tabak, Ying Wang

机构 * Department of Biomedical Signals and Systems(生物医学信号与系统系) University of Twente(埃因霍温理工大学)

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI

AI总结 本文提出CLMT模型,通过双阶段离散翻译范式解决生理信号跨模态和跨频率合成问题,显著提升临床检测精度和超分辨率性能。

详情

展开后加载摘要…

URL PDF HTML 收藏