arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11876
2508.09699 2025-08-14 cs.CV

Slot Attention-based Feature Filtering for Few-Shot Learning

Javier Rodenas, Eduardo Aguilar, Petia Radeva

机构 * AIBA, Departament de Matemàtiques & Informàtica Universitat de Barcelona(AIBA,数学与信息学院巴塞罗那大学) Departamento de Ingeniería de Sistemas y Computación Universidad Católica del Norte(系统与计算系卡罗林大学) Institute of Neuroscience, Universitat de Barcelona(神经科学研究所巴塞罗那大学)

Comments CVPR Workshop LatinX 2025

Journal ref J. Rodenas, E. Aguilar, and P. Radeva, "Slot Attention-based Feature Filtering for Few-Shot Learning," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) Workshops, 2025, pp. 30-40

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15611 2025-08-14 cs.CR

Model Poisoning Attacks to Federated Learning via Multi-Round Consistency

Yueqi Xie, Minghong Fang, Neil Zhenqiang Gong

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07755 2025-08-12 cs.CV

Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion

Minseo Kim, Minchan Kwon, Dongyeun Lee, Yunho Jeon, Junmo Kim

机构 * Korea Institute of Science and Technology(韩国科学技术院) Hanbat National University(汉拔国立大学)

Comments Accepted at CVPR 2025 workshop (AI4CC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06951 2025-08-12 cs.CV eess.IV eess.SP

SLRTP2025 Sign Language Production Challenge: Methodology, Results, and Future Work

Harry Walsh, Ed Fish, Ozge Mercanoglu Sincan, Mohamed Ilyes Lakhal, Richard Bowden, Neil Fox, Bencie Woll, Kepeng Wu, Zecheng Li, Weichao Zhao, Haodong Wang, Wengang Zhou, Houqiang Li, Shengeng Tang, Jiayi He, Xu Wang, Ruobei Zhang, Yaxiong Wang, Lechao Cheng, Meryem Tasyurek, Tugce Kiziltepe, Hacer Yalim Keles

机构 * British Machine Vision Conference(英国机器视觉会议) International Conference on Computer Vision(国际计算机视觉会议) Artificial Intelligence(人工智能) Augmented Reality(增强现实) Software Development Kit(软件开发套件) American Sign Language(美国手语) British Sign Language(英国手语) Bidirectional Long Short-Term Memory(双向长短期记忆网络) Conditional Random Field(条件随机场) Continuous Sign Language Recognition(连续手语识别) Connectionist Temporal Classification(连接主义时序分类) Content4All Deep Learning(深度学习) German Sign Language(德语手语) Swiss German Sign Language(瑞士德语手语) Dynamic Time Warping(动态时间规整) Dynamic Time Warping Mean Joint Error(动态时间规整均关节误差) Fully Connected(全连接) Feed Forward(前馈) Frames per second(帧每秒) Generative Adversarial Network(生成对抗网络) Graphics Processing Unit(图形处理单元) Gated Recurrent Unit(门控循环单元) Gloss-to-Pose Transformer(词素到姿态转换器) Gloss-to-Pose(词素到姿态) Gloss-to-Sign(词素到手语) Gloss-to-Text(词素到文本) Ground truth(真实数据) Hidden Markov Model(隐马尔可夫模型) Hand Pose Enhancer(手姿态增强器) Hard of Hearing(听力障碍) Hamburg Notation System(汉堡符号系统) Irish Sign Language(爱尔兰手语) Long Short-Term Memory(长短期记忆网络) French Sign Language(法语手语) Multi-Headed Attention(多头注意力) Motion Capture(动作捕捉) Monocular Total Capture(单目总捕捉) Mean Squared Error(均方误差) Mixture Density Network(混合密度网络) MeineDGS mDGS T Machine Translation(机器翻译) Neural Machine Translation(神经机器翻译) Natural Language Processing(自然语言处理) Non-AutoRegressive(非自回归) Noise Substitution Vector Quantization(噪声替代向量量化) PHOENIX12 RWTH-PHOENIX-Weather-2012 PHOENIX14 RWTH-PHOENIX-Weather-2014 RWTH-PHOENIX-Weather-2014 T Part Orientation Field(部分方向场) Part Of Speech(词性) Progressive Transformer(渐进转换器) Part Affinity Field(部分亲和场) Pose-to-Text Transformer(姿态到文本转换器) Pose-to-Text(姿态到文本) Pose-to-Sign(姿态到手语)

Comments 11 pages, 6 Figures, CVPR conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09688 2025-08-11 cs.CV

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

Jiho Choi, Seonho Lee, Minhyun Lee, Seungho Lee, Hyunjung Shim

机构 * KAIST, Republic of Korea(韩国科学技术院) Samsung Electronics, Republic of Korea(三星电子)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01130 2025-08-08 cs.CV

AirRoom: Objects Matter in Room Reidentification

Runmao Yao, Yi Du, Zhuoqun Chen, Haoze Zheng, Chen Wang

机构 * Spatial AI & Robotics (SAIR) Lab, University at Buffalo(空间人工智能与机器人实验室,布法罗大学)

Comments Paper accepted at CVPR 2025

Journal ref The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03069 2025-08-08 cs.CV cs.AI

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Liao Qu, Huichao Zhang, Yiheng Liu, Xu Wang, Yi Jiang, Yiming Gao, Hu Ye, Daniel K. Du, Zehuan Yuan, Xinglong Wu

机构 * ByteDance(字节跳动)

Comments CVPR 2025; Code and models: https://github.com/ByteVisionLab/TokenFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08331 2025-08-07 cs.CV

Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise

Ryan Burgert, Yuancheng Xu, Wenqi Xian, Oliver Pilarski, Pascal Clausen, Mingming He, Li Ma, Yitong Deng, Lingxiao Li, Mohsen Mousavi, Michael Ryoo, Paul Debevec, Ning Yu

机构 * Netflix Eyeline Studios(NetflixEyeline Studios) Netflix Stony Brook University(斯通布罗克大学) University of Maryland(马里兰大学) Stanford University(斯坦福大学)

Comments Accepted to CVPR'25 as Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02004 2025-08-05 cs.CV

Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention

Kyungmin Jo, Jooyeol Yun, Jaegul Choo

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01928 2025-08-05 cs.CV cs.AI cs.LG

IAUNet: Instance-Aware U-Net

Yaroslav Prytula, Illia Tsiporenko, Ali Zeynalli, Dmytro Fishman

机构 * Institute of Computer Science, University of Tartu(塔尔图大学计算机科学研究所) Ukrainian Catholic University(乌克兰天主教大学) STACC OÜ(STACC公司)

Comments Published in CVPR Workshops (CVMI), 2025. Project page/code/models/dataset: $\href{https://slavkoprytula.github.io/IAUNet/}{\text{this https URL}}$

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2025, pp. 4739-4748

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19903 2025-08-05 cs.CV

Scaling Vision Pre-Training to 4K Resolution

Baifeng Shi, Boyi Li, Han Cai, Yao Lu, Sifei Liu, Marco Pavone, Jan Kautz, Song Han, Trevor Darrell, Pavlo Molchanov, Hongxu Yin

机构 * UC Berkeley(加州大学伯克利分校) NVIDIA(英伟达)

Comments CVPR 2025. Project Page: https://nvlabs.github.io/PS3

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19853 2025-08-05 cs.CV cs.GR cs.LG

Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation

Nadav Z. Cohen, Oron Nir, Ariel Shamir

机构 * Reichman University(里奇曼大学) Microsoft Corporation(微软公司)

Comments Conference paper at CVPR 2025. Project page: https://nadavc220.github.io/conditional-balance.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11172 2025-08-04 cs.CV

TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data

Benedikt Blumenstiel, Paolo Fraccaro, Valerio Marsocci, Johannes Jakubik, Stefano Maurogiovanni, Mikolaj Czerkawski, Rocco Sedona, Gabriele Cavallaro, Thomas Brunschwiler, Juan Bernabe-Moreno, Nicolas Longépé

机构 * IBM Research – Europe(IBM欧洲研究院) European Space Agency(欧洲航天局) roman_Φ -Lab(Φ实验室) Forschungszentrum Jülich(尤利奇研究中心) University of Iceland(冰岛大学)

Comments Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10845 2025-08-04 cs.LG

Panopticon: Advancing Any-Sensor Foundation Models for Earth Observation

Leonard Waldmann, Ando Shah, Yi Wang, Nils Lehmann, Adam J. Stewart, Zhitong Xiong, Xiao Xiang Zhu, Stefan Bauer, John Chuang

Comments First two authors contributed equally. Code is available at: https://github.com/Panopticon-FM/panopticon. Accepted to CVPR 2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (2025) 2204-2214

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02254 2025-08-04 cs.CV

ProbPose: A Probabilistic Approach to 2D Human Pose Estimation

Miroslav Purkrabek, Jiri Matas

机构 * Visual Recognition Group(视觉识别组) Department of Cybernetics(cybernetics系) Faculty of Electrical Engineering(电气工程学院) Czech Technical University in Prague(布拉格捷克技术大学)

Comments Code: https://mirapurkrabek.github.io/ProbPose/

Journal ref 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01562 2025-08-04 cs.CV

Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous Circle

Miroslav Purkrabek, Jiri Matas

Comments Project Website: https://mirapurkrabek.github.io/BBox-Mask-Pose

Journal ref 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16060 2025-08-01 cs.CL

Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation

Ziqiao Ma, Jing Ding, Xuejun Zhang, Dezhi Luo, Jiahe Ding, Sihan Xu, Yuchen Huang, Run Peng, Joyce Chai

Comments COLM 2025 & CVinW @ CVPR 2025 (Spotlight). Homepage: https://vlm-reg.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18513 2025-07-30 cs.CV

LookCloser: Frequency-aware Radiance Field for Tiny-Detail Scene

Xiaoyu Zhang, Weihong Pan, Chong Bao, Xiyu Zhang, Xiaojun Xiang, Hanqing Jiang, Hujun Bao

机构 * SenseTime Research(商汤科技研究院) State Key Lab of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室)

Comments Accepted to CVPR 2025. Project page: https://coscatter.github.io/LookCloser

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2025, pages 16122-16132

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21771 2025-07-29 cs.CV

A Unified Image-Dense Annotation Generation Model for Underwater Scenes

Hongkai Lin, Dingkang Liang, Zhenghao Qi, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学)

Comments Accepted by CVPR 2025. The code is available at https://github.com/HongkLin/TIDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11921 2025-07-29 cs.CV

DeSiRe-GS: 4D Street Gaussians for Static-Dynamic Decomposition and Surface Reconstruction for Urban Driving Scenes

Chensheng Peng, Chengwei Zhang, Yixiao Wang, Chenfeng Xu, Yichen Xie, Wenzhao Zheng, Kurt Keutzer, Masayoshi Tomizuka, Wei Zhan

机构 * UC Berkeley(伯克利大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14410 2025-07-29 cs.CV cs.AI cs.LG

GLC++: Source-Free Universal Domain Adaptation through Global-Local Clustering and Contrastive Affinity Learning

Sanqing Qu, Tianpei Zou, Florian Röhrbein, Cewu Lu, Guang Chen, Dacheng Tao, Changjun Jiang

机构 * School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) School of Automotive Engineering, Tongji University(同济大学汽车工程学院) Faculty of Computer Science, Chemnitz University of Technology(化学工业大学计算机科学学院) Department of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学系) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) School of Computer Science and Technology, Key Laboratory of Embedded System and Service Computing, Ministry of Education, Tongji University(同济大学计算机科学与技术学院、教育部嵌入式系统与服务计算重点实验室)

Comments A substantial extension of the CVPR paper "Upcycling Models under Domain and Category Shift", recently accepted by IEEE-TPAMI. arXiv admin note: text overlap with arXiv:2303.07110

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12612 2025-07-28 cs.CL cs.CR

T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

Lijun Li, Zhelun Shi, Xuhao Hu, Bowen Dong, Yiran Qin, Xihui Liu, Lu Sheng, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Beihang University(北航) Harbin Institute of Technology(哈尔滨工业大学) Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)) The University of Hong Kong(香港大学)

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17585 2025-07-24 cs.CV cs.RO

From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding

Anna-Maria Halacheva, Jan-Nico Zaech, Sombit Dey, Luc Van Gool, Danda Pani Paudel

机构 * INSAIT

Comments Accepted at the OpenSUN3D Workshop, CVPR 2025. This workshop paper is not included in the official CVPR proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17032 2025-07-24 cs.CV

TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian Splatting

Jianchuan Chen, Jingchuan Hu, Gaige Wang, Zhonghua Jiang, Tiansong Zhou, Zhiwen Chen, Chengfei Lv

机构 * Alibaba Group(阿里巴巴集团)

Comments Accepted by CVPR 2025 (Highlight), project page: https://PixelAI-Team.github.io/TaoAvatar

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15265 2025-07-22 cs.CV cs.LG

Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAI

Won Jun Kim, Hyungjin Chung, Jaemin Kim, Sangmin Lee, Byeongsu Sim, Jong Chul Ye

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

Comments CVPR 2025 (poster), 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14559 2025-07-22 cs.CV

LEAD: Exploring Logit Space Evolution for Model Selection

Zixuan Hu, Xiaotong Li, Shixiang Tang, Jun Liu, Yichun Hu, Ling-Yu Duan

机构 * Peking University(北京大学) The Chinese University of Hong Kong(香港中文大学) Singapore University of Technology and Design(新加坡科技设计大学)

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13364 2025-07-21 cs.CV cs.AI

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

Siddharth Srivastava, Gaurav Sharma

Journal ref CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03557 2025-07-18 cs.CV cs.AI

Generating Synthetic Data via Augmentations for Improved Facial Resemblance in DreamBooth and InstantID

Koray Ulusan, Benjamin Kiefer

机构 * University of Tuebingen(图宾根大学) LOOKOUT University of Tuebingen(LOOKOUT 图宾根大学)

Comments Accepted to CVPR 2025 Workshop "Synthetic Data for Computer Vision Workshop", https://syndata4cv.github.io/ Revised version

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03558 2025-07-18 cs.CV

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Zehuan Huang, Yuan-Chen Guo, Xingqiao An, Yunhan Yang, Yangguang Li, Zi-Xin Zou, Ding Liang, Xihui Liu, Yan-Pei Cao, Lu Sheng

机构 * Beihang University(北京航空航天大学) VAST Tsinghua University(清华大学) The University of Hong Kong(香港大学)

Comments Project page: https://huanngzh.github.io/MIDI-Page/

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 23646 - 23657

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17237 2025-07-17 cs.CV cs.AI

Strong Baseline: Multi-UAV Tracking via YOLOv12 with BoT-SORT-ReID

Yu-Hsi Chen

机构 * The University of Melbourne(墨尔本大学)

Comments 10 pages, 5 figures, 5 tables

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) Workshops, 2025, pp. 6573-6582

详情

展开后加载摘要…

URL PDF HTML 收藏