arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11876
2502.04320 2025-07-03 cs.CV cs.LG

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Alec Helbling, Tuna Han Salih Meral, Ben Hoover, Pinar Yanardag, Duen Horng Chau

机构 * ibm(IBM研究院)

Comments Oral Presentation at ICML 2025, Best Paper Award at CVPR Workshop on Visual Concepts

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00880 2025-07-02 cs.LG cs.AI

NN-Former: Rethinking Graph Structure in Neural Architecture Representation

Ruihan Xu, Haokui Zhang, Yaowei Wang, Wei Zeng, Shiliang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) Northwestern Polytechnical University(西北工业大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

Comments Accepted to CVPR 2025. Code is avaiable at https://github.com/XuRuihan/NNFormer

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00822 2025-07-02 cs.CV

Instant Particle Size Distribution Measurement Using CNNs Trained on Synthetic Data

Yasser El Jarida, Youssef Iraqi, Loubna Mekouar

机构 * College of Computing, University Mohammed VI Polytechnic(穆拉比特大学第六理工学院)

Comments Accepted at the Synthetic Data for Computer Vision Workshop @ CVPR 2025. 10 pages, 5 figures. Code available at https://github.com/YasserElj/Synthetic-Granular-Gen

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00033 2025-07-02 cs.CV cs.AI cs.CL

Moment Sampling in Video LLMs for Long-Form Video QA

Mustafa Chasmai, Gauri Jagatap, Gouthaman KV, Grant Van Horn, Subhransu Maji, Andrea Fanelli

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Dolby Laboratories(杜比实验室)

Comments Workshop on Video Large Language Models (VidLLMs) at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02091 2025-07-02 cs.CV

Instruct-4DGS: Efficient Dynamic Scene Editing via 4D Gaussian-based Static-Dynamic Separation

Joohyun Kwon, Hanbyel Cho, Junmo Kim

机构 * DGIST KAIST(韩国国立科学技术院)

Comments Accepted to CVPR 2025. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23623 2025-07-01 cs.CV

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

Shaofei Huang, Rui Ling, Tianrui Hui, Hongyu Li, Xu Zhou, Shifeng Zhang, Si Liu, Richang Hong, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) Chinese Academy of Sciences(中国科学院) Beihang University(北航) Sangfor Technologies(深信服技术)

Comments Accepted by CVPR 2025; Code: https://github.com/spyflying/VCT_AVS; Models: https://huggingface.co/nowherespyfly/VCT_AVS

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23482 2025-07-01 cs.CV

MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting

Jun Huang, Ting Liu, Yihang Wu, Xiaochao Qu, Luoqi Liu, Xiaolin Hu

机构 * MT Lab, Meitu Inc(美图实验室,美图公司) National University of Singapore(新加坡国立大学) Department of Computer Science and Technology, BNRist, IDG/McGovern Institute for Brain Research, Tsinghua University(计算机科学与技术系,BNRist,IDG/McGovern脑科学研究院,清华大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17690 2025-07-01 cs.CV

CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model

Ziyu Yao, Xuxin Cheng, Zhiqi Huang, Lei Li

机构 * Peking University(北京大学) University of Washington(华盛顿大学) University of Copenhagen(哥本哈根大学)

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21990 2025-06-30 cs.CL cs.AI cs.LG eess.AS

Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit

Kartheek Kumar Reddy Nareddy, Sarah Ternus, Julia Niebling

机构 * Institute of Data Science(数据科学研究所) German Aerospace Center(德国航空航天中心)

Comments Computer Vision and Pattern Recognition (CVPR) 2025 Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21976 2025-06-30 cs.LG cs.AI cs.CV cs.MA cs.RO

SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model

Shuhan Tan, John Lambert, Hong Jeon, Sakshum Kulshrestha, Yijing Bai, Jing Luo, Dragomir Anguelov, Mingxing Tan, Chiyu Max Jiang

机构 * Waymo LLC(Waymo公司) UT Austin(德克萨斯大学奥斯汀分校)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10028 2025-06-27 cs.CV

Mr. DETR++: Instructive Multi-Route Training for Detection Transformers with Mixture-of-Experts

Chang-Bin Zhang, Yujie Zhong, Kai Han

机构 * The University of Hong Kong(香港大学) Meituan Inc.(美团公司)

Comments Under review. Extended version of our CVPR 2025 paper, see arXiv:2412.10028v3

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21003 2025-06-27 cs.LG

Distilling Normalizing Flows

Steven Walton, Valeriy Klyukin, Maksim Artemev, Denis Derkach, Nikita Orlov, Humphrey Shi

机构 * University of Oregon(俄勒冈大学) HSE University(俄罗斯高等经济大学) Picsart AI Research(Picsart人工智能研究所)

Comments Published in eLVM @ CVPR (https://openaccess.thecvf.com/content/CVPR2025W/eLVM/html/Walton_Distilling_Normalizing_Flows_CVPRW_2025_paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15676 2025-06-27 cs.CV

High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous Flight

Cédric Vincent, Taehyoung Kim, Henri Meeß

机构 * Télécom Paris, Institut Polytechnique de Paris(巴黎电信学院,巴黎理工大学) Fraunhofer IVI(弗劳恩霍夫IVI研究所)

Comments Accepted by CVPR2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 1461-1471

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05770 2025-06-27 cs.CV cs.AI cs.LG

Efficient Image Generation with Variadic Attention Heads

Steven Walton, Ali Hassani, Xingqian Xu, Zhangyang Wang, Humphrey Shi

机构 * University of Oregon(俄勒冈大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

Comments Published in eLVM @ CVPR (https://openaccess.thecvf.com/content/CVPR2025W/eLVM/html/Walton_Efficient_Image_Generation_with_Variadic_Attention_Heads_CVPRW_2025_paper) | Formerly named StyleNAT: Giving Each Head a New Perspective |

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20638 2025-06-26 cs.CV

Joint attitude estimation and 3D neural reconstruction of non-cooperative space objects

Clément Forray, Pauline Delporte, Nicolas Delaygue, Florence Genin, Dawa Derksen

机构 * CS Group(计算机科学组) CNES(国家空间科学中心)

Comments accepted for CVPR 2025 NFBCC workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08071 2025-06-26 cs.CV

KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points Sampling

Yu Wang, Xin Li, Shengzhao Weng, Gang Zhang, Haixiao Yue, Haocheng Feng, Junyu Han, Errui Ding

机构 * Baidu VIS(百度视觉)

Comments Accepted to CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19488 2025-06-25 cs.CV

SceneCrafter: Controllable Multi-View Driving Scene Editing

Zehao Zhu, Yuliang Zou, Chiyu Max Jiang, Bo Sun, Vincent Casser, Xiukun Huang, Jiahao Wang, Zhenpei Yang, Ruiqi Gao, Leonidas Guibas, Mingxing Tan, Dragomir Anguelov

机构 * Waymo University of Texas at Austin(德克萨斯大学奥斯汀分校) Johns Hopkins University(约翰霍普金斯大学) Google DeepMind(谷歌DeepMind)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19389 2025-06-25 cs.CV

Emergence of Text Readability in Vision Language Models

Jaeyoo Park, Sanghyuk Chun, Wonjae Kim, Sangdoo Yun, Bohyung Han

机构 * Naver AI Lab Seoul National University(首尔国立大学) Computer Vision Laboratory, ECE(计算机视觉实验室,电子与计算机工程系) IPAI

Comments EVAL-FoMo Workshop @ CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19217 2025-06-25 cs.CV cs.AI

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

Sunggu Kyung, Hyungbin Park, Jinyoung Seo, Jimin Sung, Jihyun Kim, Dongyeong Kim, Wooyoung Jo, Yoojin Nam, Sangah Park, Taehee Kwon, Sang Min Lee, Namkug Kim

机构 * Department of Biomedical Engineering, University of Ulsan College of Medicine(生物医学工程系,釜山大学医学院)

Comments 14 pages, 5 figures, submitted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19087 2025-06-25 cs.CV cs.AI

RareSpot: Spotting Small and Rare Wildlife in Aerial Imagery with Multi-Scale Consistency and Context-Aware Augmentation

Bowen Zhang, Jesse T. Boulerice, Nikhil Kuniyil, Charvi Mendiratta, Satish Kumar, Hila Shamon, B. S. Manjunath

机构 * University of California, Santa Barbara(加州大学圣芭芭拉分校) Smithsonian National Zoo and Conservation Biology Institute(史密森尼国家动物园与保护生物学研究所) Stanford University(斯坦福大学)

Comments Accepted to the CVPR 2025 Workshop on Computer Vision for Animal Behavior Tracking and Modeling (CV4Animals)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18557 2025-06-25 cs.CV

Object-aware Sound Source Localization via Audio-Visual Scene Understanding

Sung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk Kim

机构 * Kyung Hee University(庆熙大学) KAIST AI(韩国科学技术院人工智能研究所) Sungkyunkwan University(成均馆大学)

Comments Accepted at CVPR 2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 8342-8351

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18335 2025-06-25 eess.IV cs.CV

Rethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear Attention

Saad Wazir, Daeyoung Kim

机构 * School of Computing, KAIST, Republic of Korea(韩国科学技术院计算机学院)

Comments Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 30861-30871

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04293 2025-06-25 cs.CV

GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation

Weihang Li, Hongli Xu, Junwen Huang, Hyunjun Jung, Peter KT Yu, Nassir Navab, Benjamin Busam

机构 * Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) XYZ Robotics(XYZ机器人)

Comments CVPR 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03799 2025-06-25 cs.CV cs.AI

Low-power, Continuous Remote Behavioral Localization with Event Cameras

Friedhelm Hamann, Suman Ghosh, Ignacio Juarez Martinez, Tom Hart, Alex Kacelnik, Guillermo Gallego

Comments 13 pages, 8 figures, 12 tables, Project page: https://tub-rip.github.io/eventpenguins/

Journal ref IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12858 2025-06-24 cs.LG cs.CR

CDI: Copyrighted Data Identification in Diffusion Models

Jan Dubiński, Antoni Kowalczuk, Franziska Boenisch, Adam Dziedzic

机构 * Warsaw University of Technology, IDEAS NCBR(华沙理工大学,IDEAS NCBR) CISPA Helmholtz Center for Information Security(CISPA 河岸信息安全中心)

Comments Accepted at CVPR2025 (Conference on Computer Vision and Pattern Recognition) Code available at https://github.com/sprintml/copyrighted_data_identification

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17891 2025-06-24 cs.CV

Relation3D: Enhancing Relation Modeling for Point Cloud Instance Segmentation

Jiahao Lu, Jiacheng Deng

机构 * University of Science and Technology of China(中国科学技术大学)

Comments Accepted by CVPR 2025. Code: https://github.com/Howard-coder191/Relation3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17608 2025-06-24 cs.CV

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs

Nikitha SR, Aradhya Neeraj Mathur, Tarun Ram Menta, Rishabh Jain, Mausoom Sarkar

机构 * Media and Data Science Research Lab, Adobe(Adobe媒体与数据科学研究实验室)

Comments Accepted in CVPR 2025 Workshop on What's Next in Multimodal Foundational Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11748 2025-06-24 cs.CV

ILIAS: Instance-Level Image retrieval At Scale

Giorgos Kordopatis-Zilos, Vladan Stojnić, Anna Manko, Pavel Šuma, Nikolaos-Antonios Ypsilantis, Nikos Efthymiadis, Zakaria Laskar, Jiří Matas, Ondřej Chum, Giorgos Tolias

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10338 2025-06-24 cs.CV

XYScanNet: A State Space Model for Single Image Deblurring

Hanzhou Liu, Chengkai Liu, Jiacong Xu, Peng Jiang, Mi Lu

机构 * Texas A&M University(德克萨斯A&M大学) Johns Hopkins University(约翰霍普金斯大学)

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) Workshops, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18369 2025-06-24 cs.RO cs.AI cs.CV cs.SY eess.SY

G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation

Tianxing Chen, Yao Mu, Zhixuan Liang, Zanxin Chen, Shijia Peng, Qiangyu Chen, Mingkun Xu, Ruizhen Hu, Hongyuan Zhang, Xuelong Li, Ping Luo

机构 * The University of Hong Kong(香港大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Shenzhen University(深圳大学) AgileX Robotics(AgileX机器人) GDIIST HKU Shanghai Intelligent Computing Research Center(香港大学上海智能计算研究中心)

Comments Webpage: https://tianxingchen.github.io/G3Flow/, accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏