arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

1310.6424 2013-10-28 cs.LO 50%

Epistemic Logic for Communication Chains

Jeffrey Kane, Pavel Naumov

专题命中 音频语音多模态 :multi-modal(abstract)

Comments 7 pages, Contributed talk at TARK 2013 (arXiv:1310.6382) http://www.tark.org

详情

展开后加载摘要…

URL PDF HTML 收藏
1310.5089 2013-10-21 stat.ML cs.LG 50%

Kernel Multivariate Analysis Framework for Supervised Subspace Learning: A Tutorial on Linear and Kernel Multivariate Methods

Jerónimo Arenas-García, Kaare Brandt Petersen, Gustavo Camps-Valls, Lars Kai Hansen

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref IEEE Signal Processing Magazine, 30(4), 16-29, 2013

详情

展开后加载摘要…

URL PDF HTML 收藏
1207.4776 2012-07-20 cs.HC 50%

Design and User Satisfaction of Interactive Maps for Visually Impaired People

Anke Brock, Philippe Truillet, Bernard Oriola, Delphine Picard, Christophe Jouffrais

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref ICCHP 2012 (2012) 544-551

详情

展开后加载摘要…

URL PDF HTML 收藏
1206.3242 2012-06-18 cs.LG stat.ML 50%

Multi-View Learning in the Presence of View Disagreement

C. Christoudias, Raquel Urtasun, Trevor Darrell

专题命中 音频语音多模态 :audio-visual(abstract)

Comments Appears in Proceedings of the Twenty-Fourth Conference on Uncertainty in Artificial Intelligence (UAI2008)

详情

展开后加载摘要…

URL PDF HTML 收藏
1204.3726 2012-05-22 cs.DL cs.DB 50%

Proceedings of the first International Workshop On Open Data, WOD-2012

Guillaume Raschia, Martin Theobald, Ioana Manolescu

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Website of the workshop : https://sites.google.com/site/opendata2012/

详情

展开后加载摘要…

URL PDF HTML 收藏
1201.5162 2012-01-26 math.LO cs.LO 50%

A sound and complete axiomatization for Dynamic Topological Logic

David Fernández Duque

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0703070 2011-11-09 cs.HC 50%

Flexible Audio Streams

Paul Fodor

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.1454 2011-09-27 cs.HC 50%

A Prototype System for Controlling a Computer by Head Movements and Voice Commands

Anis Ismail, Abd El Salam AL Hajjar, Mohammad Hajjar

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1103.1334 2011-03-08 cs.LO 50%

Binary Sequent Calculi for Truth-invariance Entailment of Finite Many-valued Logics

Zoran Majkic

专题命中 音频语音多模态 :multi-modal(abstract)

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
0908.3362 2009-12-01 cs.HC 50%

The Function of Gesture in an Architectural Design Meeting

Willemien Visser

专题命中 音频语音多模态 :multi-modal(abstract)

Journal ref About: Designing. Analysing design meetings, Janet McDonnell and Peter Lloyd (Ed.) (2009) 269-284

详情

展开后加载摘要…

URL PDF HTML 收藏
0710.0859 2009-12-01 cs.HC 50%

Assistance orale à la recherche visuelle - étude expérimentale de l'apport d'indications spatiales à la détection de cibles

Suzanne Kieffer, Noëlle Carbonell

专题命中 音频语音多模态 :multimodal(abstract)

Comments http://www.hcirn.com/res/period/rihm.php

Journal ref Revue d'Interaction Homme-Machine 7, 1 (2006) 30 p

详情

展开后加载摘要…

URL PDF HTML 收藏
0709.0428 2009-12-01 cs.HC 50%

Oral messages improve visual search

Suzanne Kieffer, Noëlle Carbonell

专题命中 音频语音多模态 :multimodal(abstract)

Comments 4 pages

Journal ref Dans Proceedings of ACM Working Conference on Advanced Visual Interfaces - ACM Working Conference on Advanced Visual Interfaces (AVI 2006), Venezia : Italie (2006)

详情

展开后加载摘要…

URL PDF HTML 收藏
0708.3740 2009-12-01 cs.HC 50%

Plate-forme Magicien d'Oz pour l'étude de l'apport des ACAs à l'interaction

Jérôme Simonin, Marius Hategan, Noëlle Carbonell

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref Dans Actes du Second Workshop sur les Agents Conversationnels animés - Second Workshop sur les Agents Conversationnels animés, Toulouse : France (2007)

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0703071 2009-12-01 cs.OH 50%

Automatic Annotation of XHTML Pages with Audio Components

Paul Fodor

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频多模态 4772 篇

2507.14632 2026-06-17 cs.CV 版本更新 91%

BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM

BusterX++: 迈向基于MLLM的统一跨模态AI生成内容检测与解释

Haiquan Wen, Tianxiao Li, Zhenglin Huang, Yiwei He, Guangliang Cheng

机构 * University of Liverpool, UK(利物浦大学)

专题命中 视频多模态 :MLLM(title,title_cn);cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 提出统一多模态大模型BusterX++,通过纯强化学习策略实现图像与视频伪造检测的跨模态能力迁移,性能超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18558 2026-06-29 cs.CV cs.AI 版本更新 91%

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering

HiMu: 面向长视频问答的分层多模态帧选择

Dan Ben-Ami, Gabriele Serussi, Kobi Cohen, Chaim Baskin

机构 * INSIGHT Lab, Ben-Gurion University of the Negev, Israel(本-古里安内盖夫大学INSIGHT实验室,以色列) Ben-Gurion University of the Negev, Israel(本-古里安内盖夫大学,以色列)

专题命中 视频多模态 :MLLM(summary_cn,abstract);multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract)

AI总结 提出HiMu框架,通过文本LLM将查询分解为层次逻辑树,由视觉和音频专家处理原子谓词,经模糊逻辑算子组合生成连续帧满意度曲线,在16帧预算下达到帧选择方法的最优精度,并作为即插即用模块提升多种MLLM性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20818 2026-05-21 cs.CV 91%

OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026

OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026

Yisen Feng, Leigang Qu, Haoyu Zhang, Qiaohui Chu, Meng Liu, Xuemeng Song, Weili Guan, Liqiang Nie

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) National University of Singapore(新加坡国立大学) Pengcheng Laboratory(鹏城实验室) Shandong Jianzhu University(山东建筑大学) Southern University of Science and Technology(南方科技大学)

专题命中 视频多模态 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出一种基于多模态大语言模型(MLLM)的重排序框架,用于解决Ego4D事件记忆挑战2026中的自然语言查询和目标步 tracks,通过结合现有定位模型OSGNet的候选片段和MLLM的视频-语言推理能力,提升时间片段的定位精度。

Comments Champion solution for the Natural Language Queries and GoalStep tracks of the Ego4D Challenge at the CVPR EgoVis Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10819 2026-06-10 cs.CV cs.AI 新提交 90%

Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks

Earth-OneVision:将遥感多模态大语言模型扩展到更多传感器模态和任务

Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou, Yin Zhuang, Tong Zhang, Hao Wang, He Chen, Jun Li

机构 * National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing (SBIIP), Beijing Institute of Technology(北京理工大学空间智能信息处理国家重点实验室) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院) Key Laboratory of Technology in Geo-Spatial Information Processing and Application System, Chinese Academy of Sciences(中国科学院地理空间信息处理与应用系统技术重点实验室) Advanced Research Institute of Multidisciplinary Sciences, Beijing Institute of Technology(北京理工大学前沿交叉科学研究院) School of Mechatronical Engineering, Beijing Institute of Technology(北京理工大学机电学院) School of Earth and Space Sciences, Peking University(北京大学地球与空间科学学院) School of Electronics, Peking University(北京大学电子学院) School of Computer Science and Hubei Key Laboratory of Intelligent Geo-Information Processing(华中科技大学计算机科学与技术学院&湖北省智能地理信息处理重点实验室)

专题命中 视频多模态 :MLLM(summary_cn,abstract);multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出Earth-OneVision,一个2B参数的RS-MLLM,通过全粒度视觉语言对齐、空间语言同构序列化和渐进式跨模态适应机制,统一六种传感器模态和九类任务,在多个基准上达到或超越4B-72B模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26104 2026-05-26 cs.CV 90%

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding

EVIDENT: 通过实体锚定的视觉证据路由MLLM适配用于跨域视频时间定位

Geo Ahn, Jiwook Han, Youngrae Kim, Joonseok Lee, Jinwoo Choi

机构 * Kyung Hee University(庆尚大学) University of Southern California(南加州大学) Seoul National University(首尔国立大学)

专题命中 视频多模态 :MLLM(title,title_cn);分类 cs.CV

AI总结 针对视频时间定位中域迁移导致性能下降的问题,提出EVIDENT框架,通过实体瓶颈适配器、实体绑定蒸馏损失和实体到证据门控机制,利用预训练MLLM的实体注意力实现参数高效的跨域鲁棒时间定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.02662 2022-01-06 cs.CV 89%

Deep Semantic Multimodal Hashing Network for Scalable Image-Text and Video-Text Retrievals

Lu Jin, Zechao Li, Jinhui Tang

专题命中 视频多模态 :multimodal(title,abstract);image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03727 2025-10-07 cs.AI cs.CL cs.CV cs.LG 89%

Bridging the Gap Between Multimodal Foundation Models and World Models

Xuehai He

机构 * Computer Science and Engineering University of California, Santa Cruz(计算机科学与工程大学加州大学圣克ruz分校)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04751 2025-09-08 cs.IR cs.LG 89%

Multimodal Foundation Model-Driven User Interest Modeling and Behavior Analysis on Short Video Platforms

Yushang Zhao, Yike Peng, Li Zhang, Qianyi Sun, Zhihui Zhang, Yingying Zhuang

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23266 2025-08-26 cs.CV cs.AI cs.CL 89%

TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Ziyao Shangguan, Chuhan Li, Yuxuan Ding, Yanan Zheng, Yilun Zhao, Tesca Fitzgerald, Arman Cohan

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18716 2026-07-22 cs.CV 新提交 89%

Continual Video-MLLM Adaptation over Evolving Domains

在不断演变的领域上进行连续视频 - MLLM 适应

Rui Cheng, Meixing Shi, Yuxiang Cai, Jingcai Guo, Jianwei Yin, Zhi Chen

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) Hong Kong Polytechnic University(香港理工大学) The University of Southern Queensland(南昆士兰大学)

专题命中 视频多模态 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV

AI总结 研究视频多模态大语言模型对不断演变领域的适应问题,提出分布感知专家路由框架 DAER,通过保持领域隔离的轻量级专家、引入多种路由和优化机制,在域增量基准上评估,结果优于现有方法。

Comments Accepted to ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18217 2026-07-22 cs.CV 版本更新 89%

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

HOMIE:通过多模态智能增强实现以人为对象的视频个性化

Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo

机构 * Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频多模态 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.CV

AI总结 研究以人为对象的视频个性化问题,提出HOMIE框架,统一处理主体间和主体内输入设置。通过更好的MLLM集成策略、自注意力中的全局多模态引导及模态参考嵌入,在多任务中达最优性能。

Comments 28 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15689 2026-07-20 cs.CV 新提交 89%

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors

基于注意力的MLLM选择器在测试时对长视频进行高效帧选择

Yilin Wang, Xiangxi Zheng, Dongxing Mao, Linjie Li, Zhengyuan Yang, Ping Yu, Rui Yan, Yuan Yao, Alex Jinpeng Wang

机构 * ZJU(浙江大学) NJU(南京大学) CSU(中南大学) Microsoft(微软公司) NJUST(南京理工大学)

专题命中 视频多模态 :MLLM(title,title_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 研究利用MLLMs中验证选择提取层的跨模态注意力构建无训练的DAFS帧选择器,通过查询条件聚合提取帧级证据,将候选池大小和每帧令牌预算联合分配问题用动态规划解决,在Video-MME上表现出色,且无需重新训练即可跨多种模型和任务泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07850 2026-01-14 cs.MM cs.CV 88%

MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights

MLLM-VADStory: 基于领域知识的多模态大语言模型用于视频广告剧情洞察

Jasmine Yang, Poppy Zhang, Shawndra Hill

机构 * Meta

专题命中 视频多模态 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV、cs.MM

AI总结 MLLM-VADStory通过领域知识引导多模态大语言模型,系统量化和生成视频广告剧情洞察,提升视频广告创意效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09086 2024-09-17 cs.LG cs.AI cs.CV cs.DC cs.PF 88%

Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Zhenyu Ning, Jieru Zhao, Qihao Jin, Wenchao Ding, Minyi Guo

专题命中 视频多模态 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.13954 2022-08-31 cs.CV cs.SD eess.AS 88%

Video-based Cross-modal Auxiliary Network for Multimodal Sentiment Analysis

Rongfei Chen, Wenju Zhou, Yang Li, Huiyu Zhou

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16514 2026-08-18 cs.CV cs.AI cs.CL cs.HC cs.MM 新提交 88%

Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

匹配结果,不同注视:中央凹多模态大语言模型(MLLM)的搜索方式与人类的对比

Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Ulas Bagci, Alessandro Bruno

机构 * F-initiatives(F计划) Université Sorbonne Paris Nord(巴黎北索邦大学) Northwestern University(西北大学) IULM university(IULM大学)

专题命中 视频多模态 :MLLM(title_cn,summary_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究对比三款通用MLLM与人类在目标导向视觉搜索中的表现,发现模型在决策和目标获取上优于人类,但注视过程与人类不同,现有指标无法验证类人视觉,零样本模型不适用于过程层面问题。

Comments Paper accepted at 3rd HCV workshop at ECCV 2026. 12 pages main text, 16 pages supp

详情

展开后加载摘要…

URL PDF HTML 收藏