TLDW: Extreme Multimodal Summarisation of News Videos
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM
Comments Accepted at COLING-2022, disinformation, misinformation, factuality, harmfulness, fake news, propaganda, multimodality, text, images, videos, network structure, temporality
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、eess.AS
Comments CVPR2022. The final published version of the proceedings will be available on IEEE Xplore
专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments We have submitted this paper to an academic journal
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted at NAACL 2022 (Oral)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Journal ref Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR) 2022
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments 10 pages
专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments Published in the 35th Conference on Neural Information Processing Systems (NeurIPS 2021)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments NAACL 2021
专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract)
Comments Accepted by IEEE International Conference on Computer Communications 2021
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、eess.AS
专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments ECCV2020. Project page: http://movienet.site/projects/eccv20onlineperson.html
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、eess.AS
Comments To appear in the proceedings of CVPR Workshops 2020; Code: https://github.com/v-iashin/MDVC Project Page: https://v-iashin.github.io/mdvc
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Conference on Computer Vision and Pattern Recognition (CVPR 2020)
Journal ref Conference on Computer Vision and Pattern Recognition (CVPR 2020)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments IJCAI 2017
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments Resubmitted to the rebuttal for CVPR 2017 for review, 8 pages, 4 figures
博物馆视频下的资源与监管约束下的目录引导多模态归因
机构 * University of Surrey(萨里大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM
AI总结 本文提出一种目录引导的多模态归因方法,用于博物馆视频内容的自动化元数据生成,以提高档案馆发现性并满足资源和监管要求。
Comments Demo video url: https://jn00767.pages.surrey.ac.uk/catalogue-grounded-multimodal-attribution-for-museum-video/
道德愤怒影响承诺:韩国和美国YouTube上的多模态道德情感
机构 * KAIST(韩国科学技术院)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI
AI总结 本研究通过多模态道德情感分类器分析YouTube上道德愤怒对用户参与度的影响,发现其在不同文化中均能提升观看和评论等互动行为。
Comments Accepted at The Web Conference 2026. We release Korean and English multimodal moral emotion classifiers
机构 * Pennsylvania State University(宾夕法尼亚州立大学) ; Tsinghua University(清华大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 10 pages, accepted to MRAC'25: 3rd International Workshop on Multimodal and Responsible Affective Computing (ACM-MM 2025)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM;multi-modal(comments)
Comments Emotion Recognition, Deep Learning, Multi-modal, Convolutional neural network (CNN), LSTM, Situational-Knowledge, Novelty
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL;multimodal(comments)
Comments Second Grand Challenge and Workshop on Multimodal Language ACL 2020
用于高效长视频理解的证据驱动型动态视觉选择器
专题命中 视频多模态 :MLLM(summary_cn,abstract);分类 cs.CV
AI总结 本文提出基于目标MLLM内部注意力证据的动态视觉选择框架EviSelect,通过GRPO优化的随机策略实现高效长视频理解,在三个基准上性能更优,视觉token减少约50%、端到端加速3.9倍。
Comments Project Page: https://zhangbo135.github.io/EviSelect/
4DPC$^2$hat: 面向动态点云理解的失败感知自举学习
机构 * University of Science and Technology of China(中国科学技术大学)
专题命中 视频多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 提出首个针对动态点云理解的多模态大语言模型4DPC$^2$hat,通过构建大规模跨模态数据集4DPC$^2$hat-200K和引入Mamba增强的时间推理模块及失败感知自举学习策略,显著提升了动作理解与时间推理能力。
Comments Accept by ICML 2026
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments 10 pages, Accepted at ACM International Conference on Multimodal Interaction (ICMI), October 2020
Journal ref Proceedings of the 2020 International Conference on Multimodal Interaction (ICMI)
语义头专业化指导多模态大语言模型的混合视觉Transformer注意力机制
机构 * Peking University(北京大学) ; Xiaomi Corporation(小米公司) ; The University of Hong Kong(香港大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
AI总结 该研究针对多模态LLM中混合ViT注意力设计不足的问题,提出语义头专业化(SHS)概念,据此设计Ariadne注意力机制,在22项图像视频任务上性能与全注意力相当,计算量降低6.5倍。
基于多模态自我/外部中心数据采集与结构化任务知识的装配与拆卸作业中基于人工智能的工人指导
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
AI总结 该研究提出一种从专家演示中提取结构化任务知识的以数据为中心的方法,结合多模态数据推导任务表示,实现装配拆卸作业的工人指导,经案例验证其优于静态图像方法,可用于维修、培训等场景。
Comments 5 pages. Published in CIRP Annals - Manufacturing Technology
Journal ref CIRP Annals - Manufacturing Technology 75 (2026) 19-23
COMET:面向视频多模态大语言模型的对比运动增强时序推理
机构 * Peking University(北京大学) ; Nanyang Technological University(南洋理工大学) ; National University of Singapore(新加坡国立大学) ; Sun Yat-Sen University(中山大学) ; Pingan Technology(平安科技)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
AI总结 COMET是一种视频多模态大语言模型时序增强框架,通过显式时序表示等方法,在Qwen3-VL-8B等模型的动作、时序推理任务上取得性能提升,且可跨模型家族通用。
Comments Accepted at the 34th ACM International Conference on Multimedia (ACM MM 2026)