arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11877
2504.12939 2025-04-18 cs.CV cs.LG

Disentangling Polysemantic Channels in Convolutional Neural Networks

Robin Hesse, Jonas Fischer, Simone Schaub-Meyer, Stefan Roth

Comments Accepted at CVPR 2025 Workshop on Mechanistic Interpretability for Vision (MIV). Code: https://github.com/visinf/disentangle-channels

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12717 2025-04-18 cs.CV cs.AI cs.LG

Post-pre-training for Modality Alignment in Vision-Language Foundation Models

Shin'ya Yamaguchi, Dewei Feng, Sekitoshi Kanai, Kazuki Adachi, Daiki Chijiwa

Comments Accepted to CVPR 2025; Code: https://github.com/yshinya6/clip-refine

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07521 2025-04-18 cs.AI cs.MM

Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models

Yuxiang Lin, Jingdong Sun, Zhi-Qi Cheng, Jue Wang, Haomin Liang, Zebang Cheng, Yifei Dong, Jun-Yan He, Xiaojiang Peng, Xian-Sheng Hua

Comments Accepted at CVPR Workshop NEXD 2025. 21 pages, Project: https://github.com/Lum1104/EIBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00584 2025-04-18 cs.CV cs.LG

Online Video Understanding: OVBench and VideoChat-Online

Zhenpeng Huang, Xinhao Li, Jiaqi Li, Jing Wang, Xiangyu Zeng, Cheng Liang, Tao Wu, Xi Chen, Liang Li, Limin Wang

Comments CVPR 2025 Camera Ready Version. Project Page: https://videochat-online.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12284 2025-04-17 cs.CV cs.AI cs.LG

How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions

Aditya Prakash, Benjamin Lundell, Dmitry Andreychuk, David Forsyth, Saurabh Gupta, Harpreet Sawhney

Comments CVPR 2025, Project page: https://ap229997.github.io/projects/latentact

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12255 2025-04-17 cs.CV eess.IV

Human Aligned Compression for Robust Models

Samuel Räber, Andreas Plesner, Till Aczel, Roger Wattenhofer

Comments Presented at the Workshop AdvML at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12104 2025-04-17 cs.CV

Logits DeConfusion with CLIP for Few-Shot Learning

Shuo Li, Fang Liu, Zehua Hao, Xinyi Wang, Lingling Li, Xu Liu, Puhua Chen, Wenping Ma

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12021 2025-04-17 cs.CV

Action Anticipation from SoccerNet Football Video Broadcasts

Mohamad Dalal, Artur Xarles, Anthony Cioppa, Silvio Giancola, Marc Van Droogenbroeck, Bernard Ghanem, Albert Clapés, Sergio Escalera, Thomas B. Moeslund

Comments 15 pages, 14 figures. To be published in the CVSports CVPR workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12018 2025-04-17 cs.CV

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching

Xinli Yue, JianHui Sun, Junda Lu, Liangchao Yao, Fan Xia, Tianyi Wang, Fengyun Rao, Jing Lyu, Yuetang Deng

Comments Accepted to CVPR 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11879 2025-04-17 cs.CV

Learning Compatible Multi-Prize Subnetworks for Asymmetric Retrieval

Yushuai Sun, Zikun Zhou, Dongmei Jiang, Yaowei Wang, Jun Yu, Guangming Lu, Wenjie Pei

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11798 2025-04-17 cs.CV

Neighbor-Based Feature and Index Enhancement for Person Re-Identification

Chao Yuan, Tianyi Zhang, Guanglin Niu

Comments Comment: This paper has been accepted for publication in the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11786 2025-04-17 cs.CV

DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report Generation

Sang-Jun Park, Keun-Soo Heo, Dong-Hee Shin, Young-Han Son, Ji-Hye Oh, Tae-Eui Kam

Comments The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11773 2025-04-17 cs.CV

TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion

Yiran Wang, Jiaqi Li, Chaoyi Hong, Ruibo Li, Liusheng Sun, Xiao Song, Zhe Wang, Zhiguo Cao, Guosheng Lin

Comments Accepted by CVPR 2025 (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22328 2025-04-17 cs.CV cs.AI

VoteFlow: Enforcing Local Rigidity in Self-Supervised Scene Flow

Yancong Lin, Shiming Wang, Liangliang Nan, Julian Kooij, Holger Caesar

Comments CVPR 2025. Code is available at https://github.com/tudelft-iv/VoteFlow. Yancong Lin and Shiming Wang have equal contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03196 2025-04-17 cs.CV cs.HC cs.RO

SpiritSight Agent: Advanced GUI Agent with One Look

Zhiyuan Huang, Ziming Cheng, Junting Pan, Zhaohui Hou, Mingjie Zhan

Comments Paper accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04196 2025-04-17 cs.CV cs.AI

GST: Precise 3D Human Body from a Single Image with Gaussian Splatting Transformers

Lorenza Prospero, Abdullah Hamdi, Joao F. Henriques, Christian Rupprecht

Comments Camera ready for CVSports workshop at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11489 2025-04-17 cs.CV

Uncovering Branch specialization in InceptionV1 using k sparse autoencoders

Matthew Bozoukov

Comments Accepted to CVPR MIV workshop. 9 pages with an appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08012 2025-04-17 cs.CV

SRVP: Strong Recollection Video Prediction Model Using Attention-Based Spatiotemporal Correlation Fusion

Yuseon Kim, Kyongseok Park

Comments This paper has been accepted to CVPR 2025 Precognition Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11295 2025-04-16 cs.CV

Autoregressive Distillation of Diffusion Transformers

Yeongmin Kim, Sotiris Anagnostidis, Yuming Du, Edgar Schönfeld, Jonas Kohler, Markos Georgopoulos, Albert Pumarola, Ali Thabet, Artsiom Sanakoyeu

Comments CVPR 2025 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11202 2025-04-16 cs.CV eess.IV eess.SP

Focal Split: Untethered Snapshot Depth from Differential Defocus

Junjie Luo, John Mamish, Alan Fu, Thomas Concannon, Josiah Hester, Emma Alexander, Qi Guo

Comments CVPR 2025, 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11034 2025-04-16 cs.CV

Defending Against Frequency-Based Attacks with Diffusion Models

Fatemeh Amerehi, Patrick Healy

Comments Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 5th Workshop on Adversarial Machine Learning in Computer Vision: Foundation Models + X

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10857 2025-04-16 cs.RO cs.CV

ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping

Shun Iwase, Zubair Irshad, Katherine Liu, Vitor Guizilini, Robert Lee, Takuya Ikeda, Ayako Amma, Koichi Nishiwaki, Kris Kitani, Rares Ambrus, Sergey Zakharov

Comments Published at CVPR 2025, Webpage: https://sh8.io/#/zerograsp

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10738 2025-04-16 cs.CV cs.AI cs.CL cs.LG cs.RO

CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates

Ankit Kumar Shaw, Kun Jiang, Tuopu Wen, Chandan Kumar Sah, Yining Shi, Mengmeng Yang, Diange Yang, Xiaoli Lian

Comments Kun Jiang, Mengmeng Yang and Diange Yang are Corresponding Author. The main paper and supplementary material are both included here, total 23 pages (main paper is 10 pages and supplementary material is 13 pages), total 17 figures (6 figures in main paper and 11 figures in supplementary material), this paper is Accepted to CVPR WDFM-AD Workshop 2025, The code will be available at https://Ankit-Zefan.github.io/CleanMap/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10727 2025-04-16 cs.CV

Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization

Darryl Hannan, John Cooper, Dylan White, Timothy Doster, Henry Kvinge, Yijing Watkins

Comments 26 pages, CVPR MORSE Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10676 2025-04-16 cs.CV

H-MoRe: Learning Human-centric Motion Representation for Action Analysis

Zhanbo Huang, Xiaoming Liu, Yu Kong

Comments 15 pages, 14 figures, 7 tables, accepted to CVPR 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10669 2025-04-16 cs.CV cs.LG

Perturbed State Space Feature Encoders for Optical Flow with Event Cameras

Gokul Raju Govinda Raju, Nikola Zubić, Marco Cannici, Davide Scaramuzza

Comments 10 pages, 4 figures, 4 tables. Equal contribution by Gokul Raju Govinda Raju and Nikola Zubić

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Nashville, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10659 2025-04-16 cs.CV

Relation-Rich Visual Document Generator for Visual Information Extraction

Zi-Han Jiang, Chien-Wei Lin, Wei-Hua Li, Hsuan-Tung Liu, Yi-Ren Yeh, Chu-Song Chen

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10642 2025-04-16 cs.CV

SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging

Tan-Hanh Pham, Chris Ngo, Trong-Duong Bui, Minh Luu Quang, Tan-Huong Pham, Truong-Son Hy

Comments CVPR Multimodal Algorithmic Reasoning Workshop 2025 - SilVarMed

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08626 2025-04-16 cs.LG cs.AI cs.CV

Task-conditioned Ensemble of Expert Models for Continuous Learning

Renu Sharma, Debasmita Pal, Arun Ross

Comments IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Nashville, USA, June 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08548 2025-04-16 cs.GR cs.CV

COP-GEN-Beta: Unified Generative Modelling of COPernicus Imagery Thumbnails

Miguel Espinosa, Valerio Marsocci, Yuru Jia, Elliot J. Crowley, Mikolaj Czerkawski

Comments Accepted at CVPR 2025 Workshop MORSE

详情

展开后加载摘要…

URL PDF HTML 收藏