arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11877
2412.08988 2025-04-28 cs.SD cs.MM eess.AS

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

Gaoxiang Cong, Jiadong Pan, Liang Li, Yuankai Qi, Yuxin Peng, Anton van den Hengel, Jian Yang, Qingming Huang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Macquarie University(麦觉里大学) University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学) University of Adelaide(阿德莱德大学)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16718 2025-04-28 cs.CV cs.AI

Neuro-Symbolic Evaluation of Text-to-Video Models using Formal Verification

S P Sharan, Minkyu Choi, Sahil Shah, Harsh Goel, Mohammad Omama, Sandeep Chinchali

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

Journal ref Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13632 2025-04-28 cs.CV

FungiTastic: A multi-modal dataset and benchmark for image categorization

Lukas Picek, Klara Janouskova, Vojtech Cermak, Jiri Matas

机构 * University of West Bohemia & Inria, CTU in Prague(西波西米亚大学及Inria、布拉格CTU)

Comments FGVC workshop, CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02252 2025-04-28 cs.CV

StoryGPT-V: Large Language Models as Consistent Story Visualizers

Xiaoqian Shen, Mohamed Elhoseiny

机构 * KAUST(卡塞姆大学)

Comments Accepted to CVPR 2025; Project page: https://xiaoqian-shen.github.io/StoryGPT-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17788 2025-04-25 cs.CV

Dynamic Camera Poses and Where to Find Them

Chris Rockwell, Joseph Tung, Tsung-Yi Lin, Ming-Yu Liu, David F. Fouhey, Chen-Hsuan Lin

机构 * NVIDIA University of Michigan(密歇根大学) New York University(纽约大学)

Comments Accepted to CVPR 2025. Project Page: https://research.nvidia.com/labs/dir/dynpose-100k

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17695 2025-04-25 cs.CV

PICO: Reconstructing 3D People In Contact with Objects

Alpár Cseke, Shashank Tripathi, Sai Kumar Dwivedi, Arjun Lakshmipathy, Agniv Chatterjee, Michael J. Black, Dimitrios Tzionas

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) Meshcapade Carnegie Mellon University(卡内基梅隆大学) UT Austin(得克萨斯大学奥斯汀分校) University of Amsterdam(阿姆斯特丹大学)

Comments Accepted in CVPR'25. Project Page: https://pico.is.tue.mpg.de

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08585 2025-04-25 cs.CV

HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

Shehreen Azad, Vibhav Vineet, Yogesh Singh Rawat

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) Microsoft Research(微软研究院)

Comments Accepted in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17196 2025-04-25 cs.CV cs.AI

Disentangling Visual Transformers: Patch-level Interpretability for Image Classification

Guillaume Jeanneret, Loïc Simon, Frédéric Jurie

机构 * ISIR - Sorbonne University(ISIR - 索邦大学) Normandy University, ENSICAEN, UNICAEN, CNRS, GREYC(诺曼底大学、ENSICAEN、UNICAEN、CNRS、GREYC)

Comments CVPR 2025 official version. Main manuscript + supplementary

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01163 2025-04-25 cs.CV

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Jiajun Deng, Tianyu He, Li Jiang, Tianyu Wang, Feras Dayoub, Ian Reid

机构 * Australian Institute for Machine Learning, The University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Microsoft Research(微软研究院) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16684 2025-04-24 cs.CV cs.LG

SemanticSugarBeets: A Multi-Task Framework and Dataset for Inspecting Harvest and Storage Characteristics of Sugar Beets

Gerardus Croonen, Andreas Trondl, Julia Simon, Daniel Steininger

机构 * AIT Austrian Institute of Technology Center for Vision, Automation & Control(AIT奥地利技术研究所视觉、自动化与控制中心)

Comments Accepted at Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Code and dataset available at https://github.com/semanticsugarbeets/semanticsugarbeets

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13763 2025-04-24 cs.CV cs.AI

Decoding Vision Transformers: the Diffusion Steering Lens

Ryota Takatsuki, Sonia Joseph, Ippei Fujisawa, Ryota Kanai

机构 * Araya Inc.(Araya公司) AI Alignment Network(AI对齐网络) The University of Tokyo(东京大学) Mila - Quebec AI Institute(魁北克AI研究所) McGill University(麦吉尔大学)

Comments 12 pages, 17 figures. Accepted to the CVPR 2025 Workshop on Mechanistic Interpretability for Vision (MIV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01503 2025-04-24 cs.CV

Luminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve Adjustment

Ziteng Cui, Xuangeng Chu, Tatsuya Harada

机构 * The University of Tokyo(东京大学) RIKEN AIP(日本理化学研究所)

Comments CVPR 2025, project page: https://cuiziteng.github.io/Luminance_GS_web/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11515 2025-04-24 cs.CV

UltraFusion: Ultra High Dynamic Imaging using Exposure Fusion

Zixuan Chen, Yujin Wang, Xin Cai, Zhiyuan You, Zheming Lu, Fan Zhang, Shi Guo, Tianfan Xue

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Zhejiang University(浙江大学)

Comments Accepted by CVPR 2025. Project Page: https://openimaginglab.github.io/UltraFusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09466 2025-04-24 cs.CV

DEFOM-Stereo: Depth Foundation Model Based Stereo Matching

Hualie Jiang, Zhiqiang Lou, Laiyan Ding, Rui Xu, Minglang Tan, Wenjie Jiang, Rui Huang

机构 * Insta360 Research(Insta360研究院) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

Comments https://insta360-research-team.github.io/DEFOM-Stereo/

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01552 2025-04-24 cs.CV cs.RO

GFreeDet: Exploiting Gaussian Splatting and Foundation Models for Model-free Unseen Object Detection in the BOP Challenge 2024

Xingyu Liu, Gu Wang, Chengxi Li, Yingyue Li, Chenyangguang Zhang, Ziqin Huang, Xiangyang Ji

机构 * Tsinghua University(清华大学)

Comments CVPR 2025 CV4MR Workshop (citation style changed)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16030 2025-04-23 cs.CV

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

Joya Chen, Ziyun Zeng, Yiqi Lin, Wei Li, Zejun Ma, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室) ByteDance(字节跳动)

Comments CVPR 2025. If any references are missing, please contact joyachen@u.nus.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07758 2025-04-23 cs.CV eess.IV

PIDSR: Complementary Polarized Image Demosaicing and Super-Resolution

Shuangfan Zhou, Chu Zhou, Youwei Lyu, Heng Guo, Zhanyu Ma, Boxin Shi, Imari Sato

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) National Institute of Informatics(国家信息研究所) State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理国家重点实验室) National Engineering Research Center of Visual Technology, School of Computer Science, Peking University(北京大学计算机科学学院视觉技术国家工程研究中心)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15728 2025-04-23 cs.CV

SAGA: Semantic-Aware Gray color Augmentation for Visible-to-Thermal Domain Adaptation across Multi-View Drone and Ground-Based Vision Systems

Manjunath D, Aniruddh Sikdar, Prajwal Gurunath, Sumanth Udupa, Suresh Sundaram

机构 * Department of Aerospace Engineering, Indian Institute of Science(航空航天工程系,印度科学研究院) Robert Bosch Centre for Cyber Physical Systems, Indian Institute of Science(网络物理系统研究中心,印度科学研究院)

Comments Accepted at CVPR-W PBVS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15649 2025-04-23 eess.IV cs.CV

RepNet-VSR: Reparameterizable Architecture for High-Fidelity Video Super-Resolution

Biao Wu, Diankai Zhang, Shaoli Liu, Si Gao, Chengjian Zheng, Ning Wang

机构 * State Key Laboratory of Mobile Network and Mobile Multimedia Technology, ZTE, China(中国移动网络与移动多媒体技术国家重点实验室)

Comments Champion Solution for CVPR 2025 MAI VSR Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15397 2025-04-23 cs.CV

MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World

Ankit Dhiman, Manan Shah, R Venkatesh Babu

机构 * Vision and AI Lab, IISc Bangalore(视觉与人工智能实验室,班加罗尔IISc) Samsung R & D Institute India - Bangalore(三星研发研究所-印度班加罗尔)

Comments Accepted to CVPR 2025. Project Page: https://mirror-verse.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15380 2025-04-23 cs.CV

Plug-and-Play Versatile Compressed Video Enhancement

Huimin Zeng, Jiacheng Li, Zhiwei Xiong

机构 * University of Science and Technology of China(中国科学技术大学)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17820 2025-04-23 cs.CV cs.RO

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

Xinhao Liu, Jintong Li, Yicheng Jiang, Niranjan Sujay, Zhicheng Yang, Juexiao Zhang, John Abanes, Jing Zhang, Chen Feng

机构 * New York University(纽约大学)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15159 2025-04-22 cs.CV

Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration

Junyuan Deng, Xinyi Wu, Yongxing Yang, Congchao Zhu, Song Wang, Zhenyao Wu

机构 * Honor Device Co., Ltd(Honor Device公司)

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15003 2025-04-22 cs.CV

NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: KwaiSR Dataset and Study

Xin Li, Xijun Wang, Bingchen Li, Kun Yuan, Yizhen Shao, Suhang Yao, Ming Sun, Chao Zhou, Radu Timofte, Zhibo Chen

机构 * University of Science and Technology of China(中国科学技术大学) KuaiShou Technology(快手科技) University of Würzburg(乌尔姆大学)

Comments KwaiSR dataset, a new dataset for image super-resolution, used for CVPR NTIRE 2025 Challenge; CVPR 2025 workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14920 2025-04-22 cs.CV

DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding

Geng Li, Jinglin Xu, Yunzhen Zhao, Yuxin Peng

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) School of Intelligence Science and Technology, University of Science and Technology Beijing(北京科技大学智能科学与技术学院) Tencent Beijing Research(腾讯北京研究院)

Comments Accepted by CVPR 2025 (Hightlight). Project page with code: https://github.com/PKU-ICST-MIPL/DyFo_CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14875 2025-04-22 cs.CV cs.AI cs.LG

ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams

Chris Dongjoo Kim, Jihwan Moon, Sangwoo Moon, Heeseung Yun, Sihaeng Lee, Aniruddha Kembhavi, Soonyoung Lee, Gunhee Kim, Sangho Lee, Christopher Clark

机构 * Seoul National University(首尔国立大学) LG AI Research(LG人工智能研究) Allen Institute for AI(人工智能研究所)

Comments CVPR 2025 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14860 2025-04-22 cs.CV cs.AI

Bridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormer

Ziyi Liu, Yangcen Liu

机构 * University of Science and Technology Beijing(北京科技大学) Georgia Institute of Technology(佐治亚理工学院)

Comments CVPR 2025: IEEE Conference on Computer Vision and Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14687 2025-04-22 cs.CV

Seurat: From Moving Points to Depth

Seokju Cho, Jiahui Huang, Seungryong Kim, Joon-Young Lee

机构 * KAIST AI(韩国科学技术院人工智能研究所) Adobe Research(Adobe研究院)

Comments CVPR 2025 Highlight. Project page: https://seurat-cvpr.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14681 2025-04-22 cs.RO cs.AI

An LLM-enabled Multi-Agent Autonomous Mechatronics Design Framework

Zeyu Wang, Frank P. -W. Lo, Qian Chen, Yongqi Zhang, Chen Lin, Xu Chen, Zhenhua Yu, Alexander J. Thompson, Eric M. Yeatman, Benny P. L. Lo

机构 * Imperial College London(帝国理工学院伦敦分校) Univeristy of Aberdeen(阿伯丁大学)

Comments Accepted by CVPR 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14666 2025-04-22 cs.CV

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Kaihang Pan, Wang Lin, Zhongqi Yue, Tenglong Ao, Liyu Jia, Wei Zhao, Juncheng Li, Siliang Tang, Hanwang Zhang

机构 * Zhejiang University(浙江大学) Nanyang Technological University(新加坡国立大学) Peking University(北京大学) Huawei Singapore Research Center(华为新加坡研究中心)

Comments Accepted by CVPR 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏