arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11877
2506.01591 2025-06-03 cs.GR cs.CR cs.CV cs.SD eess.AS

Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation

Yuan Gan, Jiaxu Miao, Yunze Wang, Yi Yang

机构 * ReLER, CCAI, Zhejiang University(ReLER,中国人工智能学会,浙江大学) School of Cyber Science and Technology, Sun Yat-sen University(信息科学与技术学院,孙中山大学) Department of Statistics, University of Wisconsin–Madison(统计学系,威斯康星大学麦迪逊分校)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01558 2025-06-03 cs.CV

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

Yuji Wang, Haoran Xu, Yong Liu, Jiaze Li, Yansong Tang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01389 2025-06-03 cs.CV

Neural shape reconstruction from multiple views with static pattern projection

Ryo Furukawa, Kota Nishihara, Hiroshi Kawasaki

Comments 6 pages, CVPR 2025 Workshop on Neural Fields Beyond Conventional Cameras

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01304 2025-06-03 cs.CV

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost

Haiyang Mei, Pengyu Zhang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01174 2025-06-03 cs.AI

GraphPad: Inference-Time 3D Scene Graph Updates for Embodied Question Answering

Muhammad Qasim Ali, Saeejith Nair, Alexander Wong, Yuchen Cui, Yuhao Chen

机构 * University of Waterloo(滑铁卢大学) University of California, Los Angeles(加州大学洛杉矶分校)

Comments CVPR 2025 Workshop on 3D-LLM/VLA: Bridging Language, Vision and Action in 3D Environments

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01071 2025-06-03 cs.CV

Aligned Contrastive Loss for Long-Tailed Recognition

Jiali Ma, Jiequan Cui, Maeno Kazuki, Lakshmi Subramanian, Karlekar Jayashree, Sugiri Pranata, Hanwang Zhang

机构 * Panasonic R&D Center Singapore(松下新加坡研发中心) Nanyang Technological University(南洋理工大学) Panasonic Connect Co., Ltd. R&D Division(松下连接有限公司研发部)

Comments Accepted by CVPR 2025 DG-EBF Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01037 2025-06-03 cs.CV

Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution

Shijun Shi, Jing Xu, Lijing Lu, Zhihang Li, Kai Hu

机构 * Jiangnan University(江南大学) University of Science and Technology of China(中国科学技术大学) Peking University(北京大学) Chinese Academy of Sciences(中国科学院)

Comments 11 pages, 10 figures, accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00742 2025-06-03 cs.CV cs.AI

ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary

Zeqi Gu, Yin Cui, Zhaoshuo Li, Fangyin Wei, Yunhao Ge, Jinwei Gu, Ming-Yu Liu, Abe Davis, Yifan Ding

机构 * NVIDIA Cornell University(康奈尔大学)

Comments Accepted by CVPR

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23694 2025-06-03 cs.CV

DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers

Li Ren, Chen Chen, Liqiang Wang, Kien Hua

机构 * Department of Computer Science University of Central Florida(计算机科学系 佛罗里达中央大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18306 2025-06-03 cs.CV

CTRL-GS: Cascaded Temporal Residue Learning for 4D Gaussian Splatting

Karly Hou, Wanhua Li, Hanspeter Pfister

机构 * Harvard University(哈佛大学)

Comments Accepted to 4D Vision Workshop @ CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11230 2025-06-03 cs.CV cs.RO

CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image

Jingshun Huang, Haitao Lin, Tianyu Wang, Yanwei Fu, Xiangyang Xue, Yi Zhu

机构 * Fudan University(复旦大学) Huawei, Noah’s Ark Lab(华为诺亚实验室)

Comments To appear in CVPR 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00420 2025-06-03 cs.RO cs.CV

Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation

Yuanqi Yao, Siao Liu, Haoming Song, Delin Qu, Qizhi Chen, Yan Ding, Bin Zhao, Zhigang Wang, Xuelong Li, Dong Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Northwestern Polytechnical University(西北工业大学) TeleAI, China Telecom Corp Ltd(TeleAI,中国电信有限公司) INSAIT, Sofia University(INSAIT,索菲亚大学)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20418 2025-06-03 cs.CV

ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On

Ji Woo Hong, Tri Ton, Trung X. Pham, Gwanhyeong Koo, Sunjae Yoon, Chang D. Yoo

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国高级科学技术研究院)

Comments CVPR 2025, Project Page: https://jiwoohong93.github.io/ita-mdt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06029 2025-06-03 cs.CV

DiTASK: Multi-Task Fine-Tuning with Diffeomorphic Transformations

Krishna Sri Ipsit Mantri, Carola-Bibiane Schönlieb, Bruno Ribeiro, Chaim Baskin, Moshe Eliasof

机构 * Purdue University(普渡大学) University of Cambridge(剑桥大学) Ben-Gurion University of the Negev(贝内尔-阿布德大学)

Comments CVPR 2025, 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12392 2025-06-03 cs.CV cs.RO

MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors

Riku Murai, Eric Dexheimer, Andrew J. Davison

机构 * Imperial College London(帝国理工学院伦敦校区)

Comments CVPR 2025 Highlight. The first two authors contributed equally to this work. Project Page: https://edexheim.github.io/mast3r-slam/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18672 2025-06-03 cs.CV

FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models

Alice Heiman, Xiaoman Zhang, Emma Chen, Sung Eun Kim, Pranav Rajpurkar

机构 * Stanford University, USA(斯坦福大学) Harvard University, USA(哈佛大学)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17451 2025-06-03 cs.CV cs.CL

VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Lei Li, Yuancheng Wei, Zhihui Xie, Xuqing Yang, Yifan Song, Peiyi Wang, Chenxin An, Tianyu Liu, Sujian Li, Bill Yuchen Lin, Lingpeng Kong, Qi Liu

机构 * HKU(香港大学) SCUT(华南理工大学) SJTU(上海交通大学) PKU(北京大学) Allen AI(AllenAI)

Comments CVPR 2025 Camera Ready Version. Project page: https://vl-rewardbench.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03860 2025-06-03 cs.CV

MDMP: Multi-modal Diffusion for supervised Motion Predictions with uncertainty

Leo Bringer, Joey Wilson, Kira Barton, Maani Ghaffari

机构 * University of Michigan(密歇根大学)

Comments Accepted to CVPR 2025 - HuMoGen. Minor revisions made based on reviewer feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09055 2025-06-03 cs.CV

SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models

Jaerin Lee, Daniel Sungho Jung, Kanggeon Lee, Kyoung Mu Lee

机构 * ASRI, Department of ECE(ASRI电子工程系) Interdisciplinary Program in Artificial Intelligence(人工智能跨学科项目) SNU-LG AI Research Center(首尔国立大学SNU-LG人工智能研究中心)

Comments CVPR 2025 camera ready

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05984 2025-06-03 cs.CV cs.AI cs.GR cs.LG

Accurate Differential Operators for Hybrid Neural Fields

Aditya Chetan, Guandao Yang, Zichen Wang, Steve Marschner, Bharath Hariharan

机构 * Cornell University(康奈尔大学)

Comments Accepted in CVPR 2025. Project page is available at https://justachetan.github.io/hnf-derivatives/

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.00725 2025-06-03 cs.CV

CRAVES: Controlling Robotic Arm with a Vision-based Economic System

Yiming Zuo, Weichao Qiu, Lingxi Xie, Fangwei Zhong, Yizhou Wang, Alan L. Yuille

机构 * Tsinghua University(清华大学) Johns Hopkins University(约翰霍普金斯大学) Peking University(北京大学) Noah’s Ark Lab, Huawei Inc.(华为诺亚实验室) Peng Cheng Laboratory(鹏城实验室) DeepWise AI Lab(深醒人工智能实验室)

Comments 10 pages, 6 figures

Journal ref In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2019) 4214-4223

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00280 2025-06-03 cs.CR cs.CV cs.LG

3D Gaussian Splat Vulnerabilities

Matthew Hull, Haoyang Yang, Pratham Mehta, Mansi Phute, Aeree Cho, Haoran Wang, Matthew Lau, Wenke Lee, Willian T. Lunardi, Martin Andreoni, Polo Chau

机构 * Georgia Tech(佐治亚理工学院) Technology Innovation Institute(技术创新研究所)

Comments 4 pages, 4 figures, CVPR '25 Workshop on Neural Fields Beyond Conventional Cameras

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10242 2025-06-03 cs.CV cs.AI cs.LG

Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology

Oren Kraus, Kian Kenyon-Dean, Saber Saberian, Maryam Fallah, Peter McLean, Jess Leung, Vasudev Sharma, Ayla Khan, Jia Balakrishnan, Safiye Celik, Dominique Beaini, Maciej Sypetkowski, Chi Vicky Cheng, Kristen Morse, Maureen Makes, Ben Mabey, Berton Earnshaw

机构 * Recursion Valence Labs

Comments CVPR 2024 Highlight. arXiv admin note: text overlap with arXiv:2309.16064

Journal ref 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24816 2025-06-02 cs.CV

CL-LoRA: Continual Low-Rank Adaptation for Rehearsal-Free Class-Incremental Learning

Jiangpeng He, Zhihao Duan, Fengqing Zhu

机构 * Massachusetts Institute of Technology(麻省理工学院) Purdue University(普渡大学)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24703 2025-06-02 cs.CR cs.CV cs.LG

PatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial Patches

Dennis Jacob, Chong Xiang, Prateek Mittal

机构 * UC Berkeley(加州大学伯克利分校) Princeton University(普林斯顿大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24693 2025-06-02 cs.CV

Conformal Prediction for Zero-Shot Models

Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz

机构 * ÉTS Montréal(蒙特利尔ÉTS)

Comments CVPR 2025. Code: https://github.com/jusiro/CLIP-Conformal

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24690 2025-06-02 cs.CV

Learning reusable concepts across different egocentric video understanding tasks

Simone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, Tatiana Tommasi, Giuseppe Averta

机构 * Politecnico di Torino(托里诺理工学院)

Comments Extended abstract derived from arXiv:2502.02487. Presented at the Second Joint Egocentric Vision (EgoVis) Workshop (CVPR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24476 2025-06-02 cs.CV

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model

Yuting Zhang, Hao Lu, Qingyong Hu, Yin Wang, Kaishen Yuan, Xin Liu, Kaishun Wu

机构 * The Hong Kong University of Science & Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science & Technology(香港科技大学) Zhejiang University(浙江大学) Lappeenranta-Lahti University of Technology(拉佩兰塔-拉赫蒂技术大学)

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24315 2025-06-02 cs.CV

InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing

Jinlu Zhang, Yixin Chen, Zan Wang, Jie Yang, Yizhou Wang, Siyuan Huang

机构 * Center on Frontiers of Computing Studies, School of Computer Science, Peking University(1 前沿计算研究中心,计算机科学学院,北京大学) State Key Laboratory of General Artificial Intelligence, BIGAI(2 通用人工智能国家重点实验室,BIGAI) Beijing Institute of Technology(3 北京理工大学) The Chinese University of Hong Kong, Shenzhen(4 香港中文大学(深圳)) Nat’l Eng. Research Center of Visual Technology, Peking University(5 视觉技术国家工程研究中心,北京大学) Institute for AI, Peking University(6 人工智能研究院,北京大学) State Key Laboratory of General Artificial Intelligence, Peking University(7 通用人工智能国家重点实验室,北京大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17794 2025-06-02 cs.CV cs.AI

Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models

Ketan Suhaas Saichandran, Xavier Thomas, Prakhar Kaushik, Deepti Ghadiyaram

机构 * Boston University(波士顿大学) Johns Hopkins University(约翰霍普金斯大学)

Comments Accepted at CVPR 2025 workshops (AI4CC (oral) & GMCV (poster))

详情

展开后加载摘要…

URL PDF HTML 收藏