arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11877
2505.02388 2025-05-06 cs.CV cs.AI cs.LG cs.RO

MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans

Huangyue Yu, Baoxiong Jia, Yixin Chen, Yandan Yang, Puhao Li, Rongpeng Su, Jiaxin Li, Qing Li, Wei Liang, Song-Chun Zhu, Tengyu Liu, Siyuan Huang

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Beijing Institute of Technology(北京理工大学) Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02278 2025-05-06 cs.CV

Compositional Image-Text Matching and Retrieval by Grounding Entities

Madhukar Reddy Vongala, Saurabh Srivastava, Jana Košecká

机构 * Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

Comments Accepted at CVPR-W

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02236 2025-05-06 cs.CV cs.AI

Improving Physical Object State Representation in Text-to-Image Generative Systems

Tianle Chen, Chaitanya Chakka, Deepti Ghadiyaram

机构 * Boston University(波士顿大学) Runway

Comments Submitted to Synthetic Data for Computer Vision - CVPR 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02166 2025-05-06 cs.RO

CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation

Xiaoqi Li, Lingyun Xu, Mingxu Zhang, Jiaming Liu, Yan Shen, Iaroslav Ponomarenko, Jiahui Xu, Liang Heng, Siyuan Huang, Shanghang Zhang, Hao Dong

机构 * School of Computer Science, Peking University(北京大学计算机学院) PKU-Agibot Lab(北京大学AGIBOT实验室)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02148 2025-05-06 cs.CV

Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving

Alexey Nekrasov, Malcolm Burdorf, Stewart Worrall, Bastian Leibe, Julie Stephany Berrio Perez

Comments Accepted for publication at CVPR 2025. Project page: https://www.vision.rwth-aachen.de/stu-dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02071 2025-05-06 cs.CV cs.LG

Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric Learning

Can Küçüksözen, Yücel Yemez

机构 * Department of Computer Engineering, Koç University(计算机工程系,科大学院) KUIS AI Center(KUIS人工智能中心)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06457 2025-05-06 cs.CV cs.AI

Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated Learning

Yanbiao Ma, Wei Dai, Wenke Huang, Jiayi Chen

机构 * Xidian University(西安电子科技大学) Wuhan University(武汉大学)

Comments Accepted by CVPR Oral 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18674 2025-05-06 cs.CV cs.LG

Active Data Curation Effectively Distills Large-Scale Multimodal Models

Vishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans, Samuel Albanie, Federico Tombari, Yongqin Xian, Alessio Tonioni, Olivier J. Hénaff

机构 * Google(谷歌) Google DeepMind(谷歌DeepMind) Tübingen AI Center, University of Tübingen(图宾根人工智能中心,图宾根大学) University of Cambridge(剑桥大学)

Comments Accepted to IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10453 2025-05-06 cs.CV cs.GR cs.MM

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

Liu He, Yizhi Song, Hejun Huang, Pinxin Liu, Yunlong Tang, Daniel Aliaga, Xin Zhou

机构 * Purdue University(普渡大学) Baidu USA(百度美国公司) University of Rochester(罗切斯特大学)

Comments Accepted by CVPR 2025 AI4CC Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01558 2025-05-06 cs.CV

A Sensor Agnostic Domain Generalization Framework for Leveraging Geospatial Foundation Models: Enhancing Semantic Segmentation viaSynergistic Pseudo-Labeling and Generative Learning

Anan Yaghmour, Melba M. Crawford, Saurabh Prasad

机构 * University of Houston(休斯敦大学) Purdue University(普渡大学)

Comments Accepted in the 2025 CVPR Workshop on Foundation and Large Vision Models in Remote Sensing, to appear in CVPR 2025 Workshop Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01450 2025-05-06 cs.LG

Towards Film-Making Production Dialogue, Narration, Monologue Adaptive Moving Dubbing Benchmarks

Chaoyi Wang, Junjie Zheng, Zihao Chen, Shiyu Xia, Chaofan Ding, Xiaohao Zhang, Xi Tao, Xiaoming He, Xinhan Di

机构 * Shanghai Institute of Microsystem and Information Technology, CAS(上海微系统与信息工程研究所,中国科学院) AI Lab, Giant Network(巨人网络AI实验室) Fudan University(复旦大学)

Comments 6 pages, 3 figures, accepted to the AI for Content Creation workshop at CVPR 2025 in Nashville, TN

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19415 2025-05-06 cs.CV cs.AI cs.LG

AMO Sampler: Enhancing Text Rendering with Overshooting

Xixi Hu, Keyang Xu, Bo Liu, Qiang Liu, Hongliang Fei

机构 * Google(谷歌) University of Texas at Austin(德克萨斯大学奥斯汀分校)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01172 2025-05-05 cs.CV

FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component Analysis

Jiangtong Tan, Hu Yu, Jie Huang, Jie Xiao, Feng Zhao

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知联合实验室,中国科学技术大学)

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01079 2025-05-05 cs.CV eess.IV

Improving Editability in Image Generation with Layer-wise Memory

Daneul Kim, Jaeah Lee, Jaesik Park

机构 * Seoul National University(首尔国立大学)

Comments CVPR 2025. Project page : https://carpedkm.github.io/projects/improving_edit/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00998 2025-05-05 cs.CV

Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion Synthesis

Yu Hua, Weiming Liu, Gui Xu, Yaqing Hou, Yew-Soon Ong, Qiang Zhang

机构 * Nanyang Technological University(南洋理工大学) ByteDance Inc.(字节跳动公司) Dalian University(大连大学) Dalian University of Technology(大连理工大学)

Journal ref The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00975 2025-05-05 cs.CV

Generating Animated Layouts as Structured Text Representations

Yeonsang Shin, Jihwan Kim, Yumin Song, Kyungseung Lee, Hyunhee Chung, Taeyoung Na

机构 * Seoul National University(首尔国立大学) SK telecom(SK电信)

Comments AI for Content Creation (AI4CC) Workshop at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00690 2025-05-02 cs.CV cs.AI cs.RO

Towards Autonomous Micromobility through Scalable Urban Simulation

Wayne Wu, Honglin He, Chaoyuan Zhang, Jack He, Seth Z. Zhao, Ran Gong, Quanyi Li, Bolei Zhou

Comments CVPR 2025 Highlight. Project page: https://metadriverse.github.io/urban-sim/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00606 2025-05-02 cs.CV cs.LG

Dietary Intake Estimation via Continuous 3D Reconstruction of Food

Wallace Lee, YuHao Chen

机构 * University of Waterloo(滑铁卢大学)

Comments 2025 CVPR MetaFood Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00502 2025-05-02 cs.CV

Towards Scalable Human-aligned Benchmark for Text-guided Image Editing

Suho Ryu, Kihyun Kim, Eugene Baek, Dongsoo Shin, Joonseok Lee

机构 * Graduate School of Data Science, Seoul National University(数据科学研究生院,首尔国立大学)

Comments Accepted to CVPR 2025 (highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20824 2025-05-02 cs.CV eess.IV

MFSR-GAN: Multi-Frame Super-Resolution with Handheld Motion Modeling

Fadeel Sher Khan, Joshua Ebenezer, Hamid Sheikh, Seok-Jun Lee

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Samsung Research America(三星美国研究实验室)

Comments Accepted to NTIRE Workshop at CVPR 2025; 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00232 2025-05-02 cs.LG cs.AI

Scaling On-Device GPU Inference for Large Generative Models

Jiuqiang Tang, Raman Sarokin, Ekaterina Ignasheva, Grant Jensen, Lin Chen, Juhyun Lee, Andrei Kulik, Matthias Grundmann

机构 * Google LLC(谷歌有限公司) Meta Platforms, Inc(Meta平台公司)

Comments to be published in CVPR 2025 Workshop on Efficient and On-Device Generation (EDGE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21368 2025-05-01 cs.CV cs.AI

Revisiting Diffusion Autoencoder Training for Image Reconstruction Quality

Pramook Khungurn, Sukit Seripanitkarn, Phonphrm Thawatdamrongkit, Supasorn Suwajanakorn

Comments AI for Content Creation (AI4CC) Workshop at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21263 2025-05-01 cs.CV cs.LG cs.MM

Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning

Jinpeng Wang, Tianci Luo, Yaohua Zha, Yan Feng, Ruisheng Luo, Bin Chen, Tao Dai, Long Chen, Yaowei Wang, Shu-Tao Xia

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳) Shenzhen University(深圳大学) The Hong Kong University of Science and Technology(香港科学与技术大学) Research Center of Artificial Intelligence, Peng Cheng Laboratory(人工智能研究中心,鹏城实验室) Meituan, Beijing(美团,北京)

Comments Accepted by CVPR'25. 10 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11014 2025-05-01 cs.CV cs.AI

GATE3D: Generalized Attention-based Task-synergized Estimation in 3D*

Eunsoo Im, Changhyun Jee, Jung Kwon Lee

机构 * Superb AI

Comments Accepted (Poster) to the 3rd CV4MR Workshop at CVPR 2025: https://openreview.net/forum?id=00RQ8Cv3ia

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09621 2025-05-01 cs.CV

Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

Linyi Jin, Richard Tucker, Zhengqi Li, David Fouhey, Noah Snavely, Aleksander Holynski

机构 * Google DeepMind(谷歌DeepMind) University of Michigan(密歇根大学) New York University(纽约大学)

Comments CVPR 2025 Camera Ready; Data released

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05538 2025-05-01 cs.CV cs.PF

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models

Hao Cheng, Erjia Xiao, Jiayan Yang, Jiahang Cao, Qiang Zhang, Jize Zhang, Kaidi Xu, Jindong Gu, Renjing Xu

机构 * CVPR Proceedings(CVPR会议)

Comments This paper is accept by CVPR2025 (https://cvpr.thecvf.com/virtual/2025/poster/34964)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19167 2025-05-01 cs.CV cs.AI cs.RO

HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos

Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, Linguang Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, Jakob Julian Engel, Tomas Hodan

机构 * Meta Reality Labs facebookresearch(Facebook Research)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18159 2025-05-01 cs.CV

Type-R: Automatically Retouching Typos for Text-to-Image Generation

Wataru Shimoda, Naoto Inoue, Daichi Haraguchi, Hayato Mitani, Seiichi Uchida, Kota Yamaguchi

机构 * CyberAgent Kyushu University(九州大学)

Comments Accepted to CVPR 2025. Codes: https://github.com/CyberAgentAILab/Type-R

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02241 2025-05-01 cs.LG cs.AI cs.CV

Neural Redshift: Random Networks are not Random Functions

Damien Teney, Armand Nicolicioiu, Valentin Hartmann, Ehsan Abbasnejad

机构 * Idiap Research Institute(IDiap研究机构) ETH Zurich(苏黎世联邦理工学院) EPFL(洛桑联邦理工学院) University of Adelaide(阿德莱德大学)

Journal ref IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20998 2025-04-30 cs.CV cs.AI

YoChameleon: Personalized Vision and Language Generation

Thao Nguyen, Krishna Kumar Singh, Jing Shi, Trung Bui, Yong Jae Lee, Yuheng Li

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Adobe Research(Adobe研究院)

Comments CVPR 2025; Project page: https://thaoshibe.github.io/YoChameleon

详情

展开后加载摘要…

URL PDF HTML 收藏