arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11877
2505.11753 2025-05-20 cs.CV

X-Edit: Detecting and Localizing Edits in Images Altered by Text-Guided Diffusion Models

Valentina Bazyleva, Nicolo Bonettini, Gaurav Bharaj

机构 * Reality Defender

Comments CVPR (XAI4CV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11707 2025-05-20 cs.CV

Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration

Haipeng Fang, Sheng Tang, Juan Cao, Enshuo Zhang, Fan Tang, Tong-Yee Lee

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学) National Cheng-Kung University(国立成功大学)

Comments Comments: 14 pages, 14 figures. Accepted by the Proceedings of the 42nd IEEE/CVF Conference on Computer Vision and Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08851 2025-05-20 cs.LG cs.AI

Mimic In-Context Learning for Multimodal Tasks

Yuchu Jiang, Jiale Fu, Chenduo Hao, Xinting Hu, Yingzhe Peng, Xin Geng, Xu Yang

机构 * Southeast University(东南大学) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, (Southeast University), Ministry of Education(新一代人工智能技术及其跨学科应用重点实验室) Nanyang Technological University(南洋理工大学)

Comments 14 pages, 7 figures,CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08541 2025-05-20 cs.GR cs.AI cs.CV cs.RO

Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset

Zhao Dong, Ka Chen, Zhaoyang Lv, Hong-Xing Yu, Yunzhi Zhang, Cheng Zhang, Yufeng Zhu, Stephen Tian, Zhengqin Li, Geordie Moffatt, Sean Christofferson, James Fort, Xiaqing Pan, Mingfei Yan, Jiajun Wu, Carl Yuheng Ren, Richard Newcombe

机构 * Meta Reality Labs Research(Meta现实实验室) Stanford University(斯坦福大学)

Comments accepted to CVPR 2025 (Highlight). Dataset page: https://www.projectaria.com/datasets/dtc/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09694 2025-05-20 cs.CV

Omni-ID: Holistic Identity Representation Designed for Generative Tasks

Guocheng Qian, Kuan-Chieh Wang, Or Patashnik, Negin Heravi, Daniil Ostashev, Sergey Tulyakov, Daniel Cohen-Or, Kfir Aberman

机构 * Snap Research

Comments Accepted to CVPR'25. Webpage: https://snap-research.github.io/Omni-ID

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16902 2025-05-20 cs.CV cs.AI

Underwater Camouflaged Object Tracking Meets Vision-Language SAM2

Chunhui Zhang, Li Liu, Guanjie Huang, Zhipeng Zhang, Hao Wen, Xi Zhou, Shiming Ge, Yanfeng Wang

机构 * Shanghai Jiao Tong University(上海交通大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) CloudWalk Technology Co., Ltd(云walk科技有限公司) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

Comments Accepted to CVPR 2025 Workshop on CV4Animals. https://github.com/983632847/Awesome-Multimodal-Object-Tracking

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04831 2025-05-20 cs.CV

Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency

Yikai Wang, Chenjie Cao, Junqiu Yu, Ke Fan, Xiangyang Xue, Yanwei Fu

机构 * Fudan University(复旦大学) Nanyang Technological University(南洋理工大学) Alibaba DAMO Academy(阿里巴巴达摩院) Hupan Lab(虎扑实验室)

Comments CVPR 2025 Highlight. Project page: https://yikai-wang.github.io/asuka/ where full-size PDF with more qualitative results are available

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07450 2025-05-19 cs.LG cs.AI

Prototype Augmented Hypernetworks for Continual Learning

Neil De La Fuente, Maria Pilligua, Daniel Vidal, Albin Soutiff, Cecilia Curreli, Daniel Cremers, Andrey Barsky

机构 * Computer Vision Center(计算机视觉中心) TUM(技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

Comments CVPR 2025 (LatinX in CV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18711 2025-05-19 cs.CV cs.CL

Evaluating Vision-Language Models as Evaluators in Path Planning

Mohamed Aghzal, Xiang Yue, Erion Plaku, Ziyu Yao

Comments Accepted to the 2025 IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11182 2025-05-19 cs.CV cs.AI

Imputation-free and Alignment-free: Incomplete Multi-view Clustering Driven by Consensus Semantic Learning

Yuzhuo Dai, Jiaqi Jin, Zhibin Dong, Siwei Wang, Xinwang Liu, En Zhu, Xihong Yang, Xinbiao Gan, Yu Feng

机构 * National University of Defense Technology(国防科技大学) Intelligent Game and Decision Lab(智能游戏与决策实验室)

Comments The paper has been accepted by the 42nd CVPR 2025. The main text has 9 pages, including 8 figures and 4 tables. The appendix has 8 pages, with 10 figures and 6 tables. The reference list has 3 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10874 2025-05-19 cs.LG cs.AI cs.CV

MultiLink: Multi-class Structure Recovery via Agglomerative Clustering and Model Selection

Luca Magri, Filippo Leveni, Giacomo Boracchi

机构 * Politecnico di Milano(米兰理工大学)

Comments Accepted at Computer Vision and Pattern Recognition (CVPR 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10841 2025-05-19 cs.CV

RefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen Objects

Jaeguk Kim, Jaewoo Park, Keuntek Lee, Nam Ik Cho

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10810 2025-05-19 cs.CV

MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation

Gabriel Maldonado, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Vinit Katariya, Hamed Tabkhi

机构 * University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

Comments 11 pages, 5 figures, 2 tables. Presented at the CVPR 2025 Human Motion Generation (HuMoGen) Workshop. Introduces MoCLIP, a CLIP-based fine-tuning strategy for motion generation, with results on HumanML3D dataset and ablation studies

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10737 2025-05-19 cs.CV

Automated Detection of Salvin's Albatrosses: Improving Deep Learning Tools for Aerial Wildlife Surveys

Mitchell Rogers, Theo Thompson, Isla Duporge, Johannes Fischer, Klemens Pütz, Thomas Mattern, Bing Xue, Mengjie Zhang

机构 * Centre for Data Science and Artificial Intelligence(数据科学与人工智能中心) University of Otago(奥塔哥大学) Princeton University(普林斯顿大学) Marine Bycatch and Threats – Department of Conservation(海洋误捕与威胁——保护部门) Antarctic Research Trust(南极研究信托) The Tawaki Trust(塔瓦基信托) Global Penguin Society(全球企鹅协会)

Comments Accepted to the CV4Animals workshop at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10649 2025-05-19 cs.CV

Advancing Multiple Instance Learning with Continual Learning for Whole Slide Imaging

Xianrui Li, Yufei Cui, Jun Li, Antoni B. Chan

机构 * Dept. of Computer Science(计算机科学系) City University of Hong Kong(香港城市大学) Noah’s Ark Lab, Huawei Canada(华为加拿大诺亚实验室) Guangzhou Bingli Technology Co., Ltd.(广州炳利科技有限公司)

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13580 2025-05-19 cs.CV

Leveraging Automatic CAD Annotations for Supervised Learning in 3D Scene Understanding

Yuchen Rao, Stefan Ainetter, Sinisa Stekovic, Vincent Lepetit, Friedrich Fraundorfer

机构 * Inst. of Visual Computing, Graz Univ. of Technology, Austria(视觉计算研究所,格拉茨技术大学,奥地利) LIGM, École des Ponts et Chaussees, IP Paris, CNRS, France(LIGM,巴黎理工学院,IP巴黎,CNRS,法国)

Comments Project page: https://stefan-ainetter.github.io/SCANnotatepp; CVPR'25 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13303 2025-05-19 cs.CV cs.AI cs.LG

FastVLM: Efficient Vision Encoding for Vision Language Models

Pavan Kumar Anasosalu Vasu, Fartash Faghri, Chun-Liang Li, Cem Koc, Nate True, Albert Antony, Gokul Santhanam, James Gabriel, Peter Grasch, Oncel Tuzel, Hadi Pouransari

机构 * Apple(苹果公司)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10258 2025-05-16 cs.CV cs.RO

Inferring Driving Maps by Deep Learning-based Trail Map Extraction

Michael Hubbertz, Pascal Colling, Qi Han, Tobias Meisen

机构 * University of Wuppertal(乌尔姆大学) Aptiv Services Deutschland GmbH(Aptiv Services 德国公司)

Comments This paper was accepted at the CVPR WAD 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08910 2025-05-16 cs.CV cs.CL

Behind Maya: Building a Multilingual Vision Language Model

Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda, Timothy Chung, Bala Krishna S Vegesna, Abhipsha Das, Anthony Susevski, Ryan Sze-Yin Chan, S M Iftekhar Uddin, Shayekh Bin Islam, Roshan Santhosh, Snegha A, Drishti Sharma, Chen Liu, Isha Chaturvedi, Genta Indra Winata, Ashvanth. S, Snehanshu Mukherjee, Alham Fikri Aji

Comments Accepted at VLMs4ALL CVPR 2025 Workshop; corrected workshop name spelling

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02522 2025-05-16 cs.CV

Charm: The Missing Piece in ViT fine-tuning for Image Aesthetic Assessment

Fatemeh Behrad, Tinne Tuytelaars, Johan Wagemans

机构 * KU Leuven University(卢森堡大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16028 2025-05-16 cs.CV

CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images

Jungho Lee, Suhwan Cho, Taeoh Kim, Ho-Deok Jang, Minhyeok Lee, Geonho Cha, Dongyoon Wee, Dogyoon Lee, Sangyoun Lee

机构 * School of Electrical and Electronic Engineering, Yonsei University(Yonsei大学电子与电气工程学院) NAVER Cloud(NAVER云)

Comments CVPR 2025, Project Page: https://Jho-Yonsei.github.io/CoCoGaussian/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03689 2025-05-16 cs.CV cs.CL

Pose Priors from Language Models

Sanjay Subramanian, Evonne Ng, Lea Müller, Dan Klein, Shiry Ginosar, Trevor Darrell

机构 * University of California, Berkeley(加州大学伯克利分校) Google DeepMind(谷歌DeepMind) Toyota Technological Institute at Chicago(丰田技术研究所(芝加哥))

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09827 2025-05-16 cs.CV

Dyadic Mamba: Long-term Dyadic Human Motion Synthesis

Julian Tanke, Takashi Shibuya, Kengo Uchida, Koichi Saito, Yuki Mitsufuji

机构 * Sony AI(索尼人工智能) Sony Group Corporation(索尼集团)

Comments CVPR 2025 HuMoGen Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00496 2025-05-16 cs.CV

Learned Image Compression with Dictionary-based Entropy Model

Jingbo Lu, Leheng Zhang, Xingyu Zhou, Mu Li, Wen Li, Shuhang Gu

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09615 2025-05-15 cs.CV cs.SD eess.AS

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing

Yung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang, François Germain, Michael Jeffrey Jones, Moitreya Chatterjee

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国立台湾大学通信工程研究所) NVIDIA, Taiwan(台湾NVIDIA公司) Mitsubishi Electric Research Labs (MERL)(三菱电机研究实验室)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09562 2025-05-15 cs.CV

Camera-Only 3D Panoptic Scene Completion for Autonomous Driving through Differentiable Object Shapes

Nicola Marinello, Simen Cassiman, Jonas Heylen, Marc Proesmans, Luc Van Gool

机构 * KU Leuven(库勒万大学) TRACE vzw(TRACE公司) ETH Zürich(苏黎世联邦理工学院) INSAIT(INSAIT研究所)

Comments Accepted to CVPR 2025 Workshop on Autonomous Driving

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09413 2025-05-15 cs.CV

Sparse Point Cloud Patches Rendering via Splitting 2D Gaussians

Ma Changfeng, Bi Ran, Guo Jie, Wang Chongjun, Guo Yanwen

机构 * Nanjing University(南京大学) School of Software, North University of China(北中国立大学软件学院)

Comments CVPR 2025 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09393 2025-05-15 cs.GR cs.AI cs.CV

UMotion: Uncertainty-driven Human Motion Estimation from Inertial and Ultra-wideband Units

Huakun Liu, Hiroki Ota, Xin Wei, Yutaro Hirao, Monica Perusquia-Hernandez, Hideaki Uchiyama, Kiyoshi Kiyokawa

机构 * Nara Institute of Science and Technology, Japan(奈良科学技术大学,日本)

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09358 2025-05-15 cs.CV cs.LG

Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis

Bingxin Ke, Kevin Qu, Tianfu Wang, Nando Metzger, Shengyu Huang, Bo Li, Anton Obukhov, Konrad Schindler

机构 * Photogrammetry and Remote Sensing Laboratory, ETH Zürich(摄影测量与遥感实验室,苏黎世联邦理工学院)

Comments Journal extension of our CVPR 2024 paper, featuring new tasks, improved efficiency, high-resolution capabilities, and enhanced accessibility

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05507 2025-05-15 cs.RO cs.CV

AutoURDF: Unsupervised Robot Modeling from Point Cloud Frames Using Cluster Registration

Jiong Lin, Lechen Zhang, Kwansoo Lee, Jialong Ning, Judah Goldfeder, Hod Lipson

机构 * Columbia University(哥伦比亚大学) Creative Machines Lab(创意机器实验室)

Comments 16 pages

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏