arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4757 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4757 篇

2606.20919 2026-07-31 cs.CV 版本更新 74%

GIM-ENDO: A Multimodal Endoscopic Image and Video Dataset for Gastric Intestinal Metaplasia Morphology and Pathology

GIM-ENDO:用于胃肠化生形态与病理的多模态内镜图像和视频数据集

Mojgan Forootan, Mahziar Setayeshfar, Ali Darvishi, Mohammad Tashakoripour, Hamidreza Bolhasani

机构 * Gastroenterology and Liver Disease Research Center, Research Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences(沙希德·贝赫什提医科大学胃肠病与肝病研究中心,胃肠病与肝病研究所) Iran University of Medical Sciences(伊朗医科大学) Shiraz University of Medical Sciences(设拉子医科大学) Gastroenterology Department, Amiralam Hospital, Tehran University of Medical Sciences(德黑兰医科大学阿米拉拉姆医院胃肠病科) Department of Computer Engineering, Science and Research Branch, Islamic Azad University(伊斯兰阿扎德大学科学与研究分校计算机工程系)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 针对胃黏膜肠化生(GIM)公开数据集缺失的问题,构建了包含多模态内镜图像、视频及病理标注的数据集GIM-ENDO,涵盖六种主要内镜征象和亚型分级,以促进AI辅助诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16362 2026-07-21 cs.CV 新提交 74%

OmniStyle-INR: Universal and Multimodal Style Transfer for INRs

OmniStyle-INR:用于隐式神经表示的通用多模态风格迁移

Rafał Kajca, Michał Miziołek, Kornel Howil, Rafał Tobiasz, Przemysław Spurek

机构 * Jagiellonian University(雅盖隆大学) IDEAS Research Institute(IDEAS 研究所)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 研究针对不同数据模态的风格迁移问题,提出OmniStyle-INR框架,利用基于网络的连续表示作为通用域,实现了在文本提示和视觉示例引导下跨视觉模态的高质量风格迁移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16084 2026-07-07 cs.LG cs.AI 版本更新 74%

Unveiling Stochasticity: Universal Multi-modal Probabilistic Modeling for Traffic Forecasting

揭示随机性:面向交通预测的通用多模态概率建模

Weijiang Xiong, Robert Fonod, Nikolas Geroliminis

机构 * Urban Transport Systems Laboratory (LUTS), EPFL(智能交通系统实验室(LUTS),EPFL)

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

AI总结 本文提出一种通用多模态概率建模方法,通过替换输出层为GMM层,提升交通预测的不确定性建模能力,实验表明其在多种数据集上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10180 2026-07-02 cs.CV 版本更新 74%

TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval

TCMA:面向无人机跨模态文本-视频检索的文本条件多粒度对齐

Zixu Zhao, Yang Zhan, Yunhao Li, Yan Li

机构 * Carnegie Mellon University(卡内基梅隆大学) The Hong Kong Polytechnic University(香港理工大学) Wuhan University(武汉大学)

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

AI总结 针对无人机视频检索中数据集粗粒度、冗余标注问题,构建细粒度多样化标注数据集DVTMD,并提出文本条件多粒度对齐框架TCMA,实现全局-局部多层级对齐,在无人机文本-视频检索任务上达到最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12105 2026-06-11 cs.RO cs.CV cs.LG 新提交 74%

DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model

DAM-VLA: 解耦异步多模态视觉语言动作模型

Pankhuri Vanjani, Zhuoyue Li, Jakub Suliga, Moritz Reuss, Gianluca Geraci, Xinkai Jiang, Rudolf Lioutikov

机构 * Intuitive Robots Lab, Karlsruhe Institute of Technology (KIT)(直觉机器人实验室,卡尔斯鲁厄理工学院) NVIDIA(英伟达) Robotics Institute of Germany(德国机器人研究所)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 针对VLA模型同步时钟与物理交互中不同模态频率不匹配的问题,提出DAM-VLA,通过解耦各模态时间处理、维护传感器速率更新的潜在缓冲区,并利用门控交叉注意力整合高频模态,在7个真实操作任务中平均成功率提升至95.2%。

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13679 2026-06-09 cs.HC cs.CV 版本更新 74%

Toward Scalable Co-located Practical Learning: Assisting with Computer Vision and Multimodal Analytics

迈向可扩展的协同实践学习:协助计算机视觉和多模态分析

Xinyu Li, Linxuan Zhao, Yueqiao Jin, Yuchen Liu, Jin Zhou, Roberto Martinez-Maldonado, Dragan Gasevic, Lixiang Yan

机构 * Centre for Learning Analytics at Monash(墨尔本大学学习分析中心) Monash University(墨尔本大学) Department of Civil and Environmental Engineering(土木与环境工程系) School of Education(教育学院) The University of Hong Kong(香港大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 本研究评估了固定摄像头管道在重复护理模拟中的效果,通过多阶段源到目标适应提升行为检测精度,并利用行为轨迹分析提升模拟 debriefing 的可检索性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02994 2026-06-03 cs.CV 74%

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

Video-OPD:通过在线策略蒸馏实现多模态大语言模型在时序视频定位中的高效后训练

Jiaze Li, Hao Yin, Haoran Xu, Boshen Xu, Wenhui Tan, Zewen He, Jianzhong Ju, Zhenbo Luo, Jian Luan

机构 * Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 提出Video-OPD框架,利用在线策略蒸馏和教师验证分歧聚焦课程,以高效后训练多模态大语言模型进行时序视频定位,克服稀疏奖励和高计算开销问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01628 2026-06-02 q-bio.BM cs.AI 74%

Demystifying Multimodal Biomolecular Co-design With Intrinsic Geodesic Coupling

揭示具有内在测地耦合的多模态生物分子协同设计

Keyue Qiu, Xintong Wang, Zhilong Zhang, Hao Zhou, Wei-Ying Ma

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title);分类 cs.AI

AI总结 针对生物分子协同设计中模态间时间耦合被忽视的问题,提出GeoCoupling框架优化异构模态的时间耦合,在基于结构的药物设计和无条件蛋白质设计中提升物理有效性和多样性。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21458 2026-05-19 cs.CV 74%

Mining Forgery Traces from Reconstruction Error: A Weakly Supervised Framework for Multimodal Deepfake Temporal Localization

从重建误差中挖掘伪造痕迹:一种用于多模态深度伪造时间定位的弱监督框架

Midou Guo, Qilin Yin, Wei Lu, Rui Yang

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出RT-DeepLoc框架,通过重建误差识别深度伪造,利用MAE学习真实数据的时空模式,结合不对称视频对比损失提升定位精度,实验表明在大规模数据集上达到弱监督时间伪造定位的最新水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05959 2026-04-29 cs.CV cs.LG 74%

Multi-Modal Landslide Detection from Sentinel-1 SAR and Sentinel-2 Optical Imagery Using Multi-Encoder Vision Transformers and Ensemble Learning

基于Sentinel-1 SAR和Sentinel-2光学影像的多模态滑坡检测:使用多编码器视觉Transformer和集成学习

Ioannis Nasios

机构 * NodalPoint

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 本文提出融合Sentinel-2光学影像与Sentinel-1 SAR数据的多模型框架,利用多编码器视觉Transformer和集成学习提升滑坡检测精度,实现91.9的F1分数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14630 2026-04-17 cs.CV cs.LG 74%

CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation

CMTM:跨模态令牌调制用于无监督视频对象分割

Inseok Jeon, Suhwan Cho, Minhyeok Lee, Seunghoon Lee, Minseok Kang, Jungho Lee, Chaewon Park, Donghyeong Kim, Sangyoun Lee

机构 * Yonsei University, Republic of Korea(延世大学,韩国)

专题命中 视频多模态 :cross-modal(title);分类 cs.CV

AI总结 本文提出CMTM方法,通过跨模态令牌调制增强外观与运动提示的交互,实现高效的多模态信息传播,取得最佳性能。

Comments 6 pages, 5 figures. Accepted to IEEE ICIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16806 2026-04-15 cs.CV 74%

MedGS: Gaussian Splatting for Multi-Modal 3D Medical Imaging

MedGS:用于多模态3D医学影像的高斯点散射

Kacper Marzol, Ignacy Kolton, Weronika Smolak-Dyżewska, Joanna Kaleta, Żaneta Świderska-Chadaj, Marcin Mazur, Mirosław Dziekiewicz, Tomasz Markiewicz, Przemysław Spurek

机构 * Jagiellonian University, Poland(雅盖隆大学,波兰) Warsaw University of Technology, Poland(华沙理工大学,波兰) IDEAS Research Institute, Poland(IDEAS研究所,波兰) Military Institute of Medicine, Poland(军事医学研究所,波兰)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 本文提出MedGS,一种利用内镜成像独特性质的3D重建框架,通过物理逼真重照明模型提升重建质量,实现更准确的临床应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02093 2026-04-03 cs.CV 74%

GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding

GroundVTS:多模态大语言模型中的视频时间定位视觉标记采样

Rong Fan, Kaiyan Xiao, Minghao Zhu, Liuyi Wang, Kai Dai, Zhao Yang

机构 * Newcapec AI Research(新开普人工智能研究院) Fudan University(复旦大学) Tongji University(同济大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 本文提出GroundVTS,通过细粒度查询引导机制筛选视觉标记,提升视频时间定位的时空信息保留与时间连贯性,实验表明其在三个标准基准上优于现有方法。

Comments Published as a conference paper at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01712 2026-04-03 cs.LG cs.AI eess.SP physics.comp-ph 74%

Transformer self-attention encoder-decoder with multimodal deep learning for response time series forecasting and digital twin support in wind structural health monitoring

基于多模态深度学习的Transformer自注意力编码器-解码器:用于响应时间序列预测和风力结构健康监测的数字孪生支持

Feiyu Zhou, Marios Impraimakis

机构 * Department of Mechanical Engineering, University of Bath(巴斯大学机械工程系) Department of Civil Engineering, Zhejiang University(浙江大学土木工程系)

专题命中 视频多模态 :multimodal(title);分类 cs.AI

AI总结 本文提出一种基于Transformer的多模态深度学习方法,用于风力结构响应时间序列预测和数字孪生支持,通过捕捉系统时间特性提升结构健康监测的准确性与鲁棒性。

Comments 21 pages, 22 figures, 9 tables. This version corresponds to the published article in Computers & Structures. https://doi.org/10.1016/j.compstruc.2026.108216

Journal ref Computers and Structures 326 (2026) 108216

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16997 2026-03-31 cs.CV cs.LG cs.RO 74%

Resolving Spatio-Temporal Entanglement in Video Prediction via Multi-Modal Attention

通过多模态注意力解决视频预测中的时空纠缠

Shreyam Gupta, P. Agrawal, Priyam Gupta

机构 * Indian Institute of Technology (BHU), Varanasi(印度理工学院(巴纳拉斯印度教大学),瓦拉纳西) University of Colorado, Boulder(科罗拉多大学博尔德分校) Erasmus+, Intelligent Field Robotic Systems (IFRoS), University of Girona(伊拉斯谟+,智能现场机器人系统(IFRoS),赫罗纳大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 本文提出MAUCell架构,结合GAN与层级处理策略及三种注意力机制,解决RNN在长时间序列中的局限,提升视频预测的时空一致性与实时性。

Comments 11 pages, 3 figures, 5 tables, and 3 Algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18990 2026-03-13 cs.CV 74%

IDSelect: A RL-Based Cost-Aware Selection Agent for Video-based Multi-Modal Person Recognition

IDSelect: 一种基于强化学习的成本感知选择代理用于基于视频的多模态人识别

Yuyang Ji, Yixuan Shen, Kien Nguyen, Lifeng Zhou, Feng Liu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 IDSelect 通过强化学习选择最优模型提升多模态人识别的效率与准确度,实验显示其在计算资源上显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01284 2026-03-03 cs.CV 74%

FoSS: Modeling Long Range Dependencies and Multimodal Uncertainty in Trajectory Prediction via Fourier State Space Integration

FoSS:通过傅里叶状态空间整合建模长距离依赖性和多模态不确定性在轨迹预测中

Yizhou Huang, Gengze Jiang, Yihua Cheng, Kezhi Wang

机构 * Brunel University of London(伦敦布鲁内尔大学) University of Birmingham(伯明翰大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 FoSS通过结合频域和时域建模,有效提升轨迹预测的准确性和效率,减少计算与参数消耗,实现多模态不确定性建模。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22923 2026-02-27 cs.CV cs.RO 74%

WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents

WaterVideoQA: 以ASV为中心的感知与符合规则的推理 via 多模态智能体

Runwei Guan, Shaofeng Liang, Ningwei Ouyang, Weichen Fei, Shanliang Yao, Wei Dai, Chenhao Ge, Penglei Sun, Xiaohui Zhu, Tao Huang, Ryan Wen Liu, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)人工智能研究所) Hubei Key Laboratory of Inland Shipping Technology (Wuhan University of Technology)(湖北内河航运技术重点实验室(武汉理工大学)) School of Navigation, Wuhan University of Technology(武汉理工大学航海学院) School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学先进科技学院) School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) School of Information Engineering, Yancheng Institute of Technology(盐城职业技术学院信息工程学院) School of Engineering, Stanford University(斯坦福大学工程学院) Centre for AI and Data Science Innovation and the School of Science and Engineering, James Cook University(詹姆斯库克大学人工智能与数据科学创新中心及科学与工程学院)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 WaterVideoQA通过多模态智能体系统,实现ASV在复杂水域环境中的感知与规则合规推理,提升自主航行的安全性和精确性。

Comments 11 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08683 2026-02-27 cs.CV 74%

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

OneVision-Encoder: 编码对齐的稀疏性作为多模态智能的基础原则

Feilong Tang, Xiang An, Yunyao Yan, Yin Xie, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Chunyuan Li, Shikun Feng, Changrui Chen, Huajie Tan, Ming Hu, Manyuan Zhang, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng

机构 * Glint Lab(Glint实验室) AIM for Health Lab(健康人工智能实验室) MVP Lab(MVP实验室)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 OneVision-Encoder通过编码对齐的稀疏性原则,实现高效的多模态视觉理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17785 2026-02-23 cs.CV 74%

Multi-Modal Monocular Endoscopic Depth and Pose Estimation with Edge-Guided Self-Supervision

多模态单目内窥镜深度与姿态估计与边缘引导自监督学习

Xinwei Ju, Rema Daher, Danail Stoyanov, Sophia Bano, Francisco Vasconcelos

机构 * UCL Hawkes Institute, Department of Computer Science(伦敦大学学院UCL霍克斯研究所、计算机科学系)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 PRISM通过边缘引导自监督学习,结合解剖学和照明先验知识,提升内窥镜单目深度与姿态估计性能。

Comments 14 pages, 6 figures; early accepted by IPCAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00132 2026-02-03 cs.CV 74%

Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time Adaptation

去除伪装,连接领域:通过测试时适应检测转移多模态仇恨视频

Jiao Li, Jian Lang, Xikai Tang, Wenzheng Shu, Ting Zhong, Qiang Gao, Yong Wang, Leiting Chen, Fan Zhou

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 SCANNER通过测试时适应框架,利用仇恨内容中稳定的内核连接源与目标领域,有效应对多模态仇恨视频检测中的语义漂移问题。

Comments Accepted by AAAI2026 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22661 2026-01-22 cs.IR cs.AI 74%

Next Point-of-interest (POI) Recommendation Model Based on Multi-modal Spatio-temporal Context Feature Embedding

基于多模态时空上下文特征嵌入的下一个兴趣点(POI)推荐模型

Lingyu Zhang, Pengfei Xu, Rui Ban, Zhenchao Zhang, Songtao Liu, Yan Wang, Yunhai Wang

机构 * Institute of Trustworthy Autonomous Systems, Southern University of Science and Technology(SUSTech)(可信自主系统研究院,南方科技大学) School of Information Sciences and Technology, Northwest University(信息科学与技术学院,西北大学) China Information Technology Designing & Consulting Institute Co., Ltd.(中国信息科技设计与咨询研究院有限公司) Renmin University of China(中国人民大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

AI总结 本文提出基于多模态时空上下文特征嵌入的POI推荐模型,通过双流时空注意力机制有效区分长期习惯与短期意图,提升个性化出行预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19850 2026-01-09 cs.CV 74%

FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs

FALCONEye: 在一小时内视频中寻找答案并定位内容的多模态大语言模型

Carlos Plou, Cesar Borja, Ruben Martinez-Cantin, Ana C. Murillo

机构 * DIIS-I3A, University of Zaragoza(DIIS-I3A,西班牙阿利坎特大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 FALCONEye通过多模态大语言模型在小时长视频中高效定位内容并回答问题,超越现有模型性能并显著降低推理成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23361 2026-01-01 cs.CV 74%

OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions

OmniVCus: 基于多模态控制条件的前馈主体驱动视频定制

Yuanhao Cai, He Zhang, Xi Chen, Jinbo Xing, Yiwei Hu, Yuqian Zhou, Kai Zhang, Zhifei Zhang, Soo Ye Kim, Tianyu Wang, Yulun Zhang, Xiaokang Yang, Zhe Lin, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) Adobe Research(Adobe研究) The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

AI总结 OmniVCus通过多模态控制条件和改进的嵌入机制实现高效的多主体视频定制。

Comments NeurIPS 2025; A data construction pipeline and a diffusion Transformer framework for controllable subject-driven video customization

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21693 2025-12-01 cs.MM 74%

Designing a Multimodal Viewer for Piano Performance Analysis -- a Pedagogy-First Approach

为钢琴表演分析设计多模态查看器——一种以教学为导向的方法

Joonhyung Bae, Hyeyoon Cho, Kirak Kim, Dawon Park, Taegyun Kwon, Yoon-Seok Choi, Hyeon Hur, Shigeru Kai, Yohei Wada, Satoshi Obata, Akira Maezawa, Jaebum Park, Jonghwa Park, Juhan Nam

专题命中 视频多模态 :multimodal(title);分类 cs.MM

AI总结 本研究提出了一种以教学为导向的多模态查看器,通过整合视频、动作捕捉和乐谱,为钢琴教学提供具体的视觉反馈,以提高教学指导的清晰度和有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16524 2025-11-21 cs.CV 74%

BoxingVI: A Multi-Modal Benchmark for Boxing Action Recognition and Localization

BoxingVI:一种多模态基准,用于拳击动作识别与定位

Rahul Kumar, Vipul Baghel, Sudhanshu Singh, Bikash Kumar Badatya, Shivam Yadav, Babji Srinivasan, Ravi Hegde

机构 * Indian Institute of Technology Gandhinagar(印度理工学院甘地纳加尔) Indian Institute of Technology Madras(印度理工学院马德拉斯) Dr. A. P. J. Abdul Kalam Technical University(阿卜杜勒·卡拉姆技术大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 BoxingVI提供了一个多模态数据集,用于拳击动作识别与定位,旨在促进低资源环境下的实时视觉动作识别研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15342 2025-11-20 cs.HC cs.AI 74%

Reflexive Evidence-Based Multimodal Learning for Clean Energy Transitions: Causal Insights on Cooking Fuel Access, Urbanization, and Carbon Emissions

Shan Shan

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01562 2025-11-17 cs.CV 74%

Adaptive LiDAR Scanning: Harnessing Temporal Cues for Efficient 3D Object Detection via Multi-Modal Fusion

Sara Shoouri, Morteza Tavakoli Taba, Hun-Seok Kim

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01481 2025-10-28 cs.CV cs.LG 74%

VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding

Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin, Hongyang Du, Fuxiao Liu, Tianyi Zhou, Dinesh Manocha, Jordan Lee Boyd-Graber

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06509 2025-10-13 cs.CV 74%

From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding

Shih-Yao Lin, Sibendu Paul, Caren Chen

机构 * Amazon Prime Video(亚马逊Prime视频)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏