VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
机构 * University of Maryland, College Park(马里兰大学学院公园分校)
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Journal ref NeurIPS 2025
AI 大模型
视频理解、视频生成、视频语言模型和时序视觉推理。
机构 * University of Maryland, College Park(马里兰大学学院公园分校)
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Journal ref NeurIPS 2025
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Zhejiang University(浙江大学) ; The University of Tokyo(东京大学) ; Fudan University(复旦大学) ; Nanjing University(南京大学)
专题命中 视频理解 :video reasoning(abstract);分类 cs.CV
Comments Accepted at NeurIPS 2025
机构 * Southeast University(东南大学) ; Monash University(墨尔本大学) ; Xiaohongshu Inc.(小红书公司) ; University of Southern California(南加州大学) ; Fudan University(复旦大学)
专题命中 视频理解 :video reasoning(abstract);分类 cs.CV
机构 * University of Science and Technology of China(中国科学技术大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Beihang University(北京航空航天大学) ; Shanghai Jiao Tong University(上海交通大学) ; Zhejiang University(浙江大学) ; State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) ; Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学) ; HKUST(香港科技大学) ; Xmax.AI Ltd(Xmax人工智能有限公司) ; HIT (SZ)(哈尔滨工业大学(深圳))
专题命中 视频生成 :video generation(title,abstract);text-to-video(title,abstract);分类 cs.CV
机构 * Shanghai Jiao Tong University, Tencent Youtu Lab(上海交通大学,腾讯云图实验室) ; Tencent Youtu Lab(腾讯云图实验室) ; Shanghai Jiao Tong University(上海交通大学) ; Tencent(腾讯)
专题命中 视频生成 :text-to-video(title,abstract);video generation(title);分类 cs.CV
Comments ACM Multimedia 2025; code URL: https://github.com/rain152/IPVG
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; University of Chicago(芝加哥大学) ; Stanford University(斯坦福大学)
专题命中 视频生成 :video generation(title,abstract);long video(abstract);分类 cs.CV
机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院) ; Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 视频生成 :video generation(title,abstract);分类 cs.CV、cs.MM
Comments Accepted by ICCV 2025
机构 * MMLab, CUHK(CUHK媒体实验室) ; Tsinghua University(清华大学) ; Kling Team, Kuaishou Technology(快手科技 Kling 团队) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 视频生成 :video generation(title,abstract);分类 cs.CV
机构 * FAIR, Meta Superintelligence Labs(FAIR、Meta超智能实验室) ; University of Oxford(牛津大学) ; Mila - Québec AI Institute(魁北克人工智能研究所) ; Université de Montréal(蒙特利尔大学) ; Columbia University(哥伦比亚大学) ; McGill University(麦吉尔大学) ; Canada CIFAR AI Chair(加拿大CIFAR人工智能 chair)
专题命中 视频生成 :video generation(title);分类 cs.CV
Comments 2 pages
Journal ref Winning entry of the ICCV 2025 Physics IQ Challenge
机构 * Adobe Research(Adobe研究机构) ; University of Michigan(密歇根大学) ; UNC Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 视频生成 :video diffusion(abstract);分类 cs.CV
Comments ICCV 2025; First three authors contributed equally. Project page: https://veggie-gen.github.io/
机构 * Huawei Technologies Canada(华为技术加拿大)
专题命中 视频扩散模型 :video generation(abstract)
Comments 16 pages, 11 figure, 2 tables, accepted at Neurips 2025
机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) ; Nanyang Technological University(南洋理工大学)
专题命中 视频问答 :video reasoning(abstract);分类 cs.CV
Comments EMNLP 2025 Findings; The first two authors contributed equally; Github link: https://github.com/Yui010206/MEXA
机构 * Shanghai Jiao Tong University(上海交通大学) ; Monash University(墨尔本大学) ; Zhongguancun Academy(中关村academy) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 动作与事件理解 :video reasoning(title);video understanding(abstract);long video(abstract);分类 cs.CV
Comments Accepted by IEEE TPAMI (IEEE Transactions on Pattern Analysis and Machine Intelligence). arXiv admin note: substantial text overlap with arXiv:2409.17647
机构 * School of Computer Science and Electronic Engineering, University of Surrey(Surrey大学计算机科学与电子工程学院) ; School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院) ; Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey大学视觉、语音和信号处理中心)
专题命中 动作与事件理解 :video generation(abstract);long video(abstract);分类 cs.CV
机构 * Department of Computer Science(计算机科学系) ; Delft University of Technology(代尔夫特理工大学) ; LatentWorlds AI ; National Policelab AI & Model-Driven Decisions Lab(国家警务实验室AI与模型驱动决策实验室)
专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV
Comments Accepted as poster in the NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI