arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11876
2505.21755 2025-06-24 cs.CV cs.AI cs.CL cs.LG

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering

Chengyue Huang, Brisa Maneechotesuwan, Shivang Chopra, Zsolt Kira

机构 * Georgia Institute of Technology(佐治亚理工学院)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16589 2025-06-23 cs.CV cs.AI cs.PF stat.ML

Spatially-Aware Evaluation of Segmentation Uncertainty

Tal Zeevi, Eléonore V. Lieffrig, Lawrence H. Staib, John A. Onofrey

机构 * Yale University(耶鲁大学)

Comments Presented at the 4th Workshop on Uncertainty Quantification for Computer Vision (CVPR 2025), June 11, 2025. This version is not included in the official proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13183 2025-06-23 cs.CV

Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures

Guoxing Sun, Rishabh Dabral, Heming Zhu, Pascal Fua, Christian Theobalt, Marc Habermann

机构 * Max Planck Institute for Informatics(马克斯·普朗克研究所信息学研究所) VIA Research Center(VIA研究中心) EPFL(瑞士联邦理工学院)

Comments Accepted at CVPR 2025, Project page: https://vcai.mpi-inf.mpg.de/projects/DUT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15720 2025-06-23 cs.LG cs.CV

Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learning

Juntae Lee, Munawar Hayat, Sungrack Yun

机构 * Qualcomm AI Research(高通人工智能研究)

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12919 2025-06-23 cs.LG cond-mat.mtrl-sci

Bridging Text and Crystal Structures: Literature-driven Contrastive Learning for Materials Science

Yuta Suzuki, Tatsunori Taniai, Ryo Igarashi, Kotaro Saito, Naoya Chiba, Yoshitaka Ushiku, Kanta Ono

机构 * Toyota Motor Corporation(丰田汽车公司) OMRON SINIC X Corporation(OMRON SINIC X公司) Randeft, Inc.(Randeft公司) The University of Osaka(大阪大学)

Comments 19 pages, 8 figures. Accepted to Machine Learning: Science and Technology (2025). Preliminary versions appeared at NeurIPS 2024 AI4Mat and CVPR 2025 MM4Mat workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20499 2025-06-19 cs.LG

Data Distributional Properties As Inductive Bias for Systematic Generalization

Felipe del Rio, Alain Raymond-Saez, Daniel Florea, Rodrigo Toro Icarte, Julio Hurtado, Cristian B. Calderon, Alvaro Soto

机构 * Pontificia Universidad Católica de Chile(天主教智利大学) University of Warwick(沃里克大学) CENIA

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14808 2025-06-19 cs.LG

PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models

Jenny Schmalfuss, Nadine Chang, Vibashan VS, Maying Shen, Andres Bruhn, Jose M. Alvarez

机构 * University of Stuttgart(斯图加特大学) NVIDIA(NVIDIA公司) Johns Hopkins University(约翰霍普金斯大学)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14263 2025-06-18 cs.LG math.OC

Towards Robust Learning to Optimize with Theoretical Guarantees

Qingyu Song, Wei Lin, Juncheng Wang, Hong Xu

Comments Published in CVPR 2024, 55 pages, 17 figures, this version fixed some typo

Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, 17297-17306

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13780 2025-06-18 cs.CV cs.AI cs.CY cs.LG

Hidden Bias in the Machine: Stereotypes in Text-to-Image Models

Sedat Porikli, Vedat Porikli

机构 * Canyon Crest Academy(山崖顶峰学院)

Comments Equal contribution by both authors, Published at CVPR 2025 Workshop on Experimental Model Auditing via Controllable Synthesis (EMACS) and Workshop on Demographic Diversity in Computer Vision (DemoDiv)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19834 2025-06-18 cs.LG cs.CV cs.MM

Knowledge Bridger: Towards Training-free Missing Modality Completion

Guanzhou Ke, Shengfeng He, Xiao Li Wang, Bo Wang, Guoqing Chao, Yuanyang Zhang, Yi Xie, HeXing Su

机构 * Beijing Jiaotong University(北京交通大学) Singapore Management University(新加坡国立大学) Southeast University(东南大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Harbin Institute of Technology(哈尔滨工业大学) Nanjing University of Science and Technology(南京理工大学) South China University of Technology(华南理工大学) Xiamen Institute of Technology(厦门理工学院)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10794 2025-06-18 cs.CV

Distraction is All You Need for Multimodal Large Language Model Jailbreaking

Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao, Xin Lin, Tao Li, Kanghua Mo, Changyu Dong

机构 * Guangzhou University(广州大学) Shanghai Jiao Tong University(上海交通大学) Australian Institute for Machine Learning, The University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学)

Comments CVPR 2025 highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18562 2025-06-18 cs.RO cs.CV cs.LG

DexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous Manipulation

Zhixuan Liang, Yao Mu, Yixiao Wang, Tianxing Chen, Wenqi Shao, Wei Zhan, Masayoshi Tomizuka, Ping Luo, Mingyu Ding

Comments Accepted by CVPR 2025. Camera ready version. Previous DexDiffuser. Project page: https://dexdiffuser.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12992 2025-06-17 cs.CV

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models

Xinyi Zhao, Congjing Zhang, Pei Guo, Wei Li, Lin Chen, Chaoyue Zhao, Shuai Huang

机构 * University of Washington(华盛顿大学) Wyze Labs, Inc.(Wyze实验室)

Comments CVPR 2025 Workshop: VAND 3.0 - Visual Anomaly and Novelty Detection

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12716 2025-06-17 cs.CV

Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors

Wen-Hsuan Chu, Lei Ke, Jianmeng Liu, Mingxiao Huo, Pavel Tokmakov, Katerina Fragkiadaki

机构 * Carnegie Mellon University(卡内基梅隆大学) Toyota Research Institute(丰田研究中心)

Comments This is an updated and extended version of our CVPR paper "Robust Multi-Object 4D Generation in Complex Video Scenarios"

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12585 2025-06-17 cs.CV

DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification

Darryl Ho, Samuel Madden

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

Comments Accepted to CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition), main conference, poster presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08210 2025-06-17 cs.CV cs.AI cs.CL cs.LG

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation

Andrew Z. Wang, Songwei Ge, Tero Karras, Ming-Yu Liu, Yogesh Balaji

机构 * University of Maryland(马里兰大学) NVIDIA(英伟达)

Comments CVPR 2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 28575-28585

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24402 2025-06-17 cs.CV

Leveraging Intermediate Features of Vision Transformer for Face Anti-Spoofing

Mika Feng, Koichi Ito, Takafumi Aoki, Tetsushi Ohki, Masakatsu Nishigaki

机构 * Graduate School of Information Sciences, Tohoku University, Japan(东北大学信息科学研究生院) Faculty of Informatics, Shizuoka University, Japan(静冈大学信息学系)

Comments 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10622 2025-06-17 cs.LG cs.AI cs.CL cs.CV

Transformers without Normalization

Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun, Zhuang Liu

机构 * FAIR, Meta(FAIR与Meta公司) New York University(纽约大学) MIT(麻省理工学院) Princeton University(普林斯顿大学)

Comments CVPR 2025; Project page: https://jiachenzhu.github.io/DyT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02357 2025-06-17 cs.CV

Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content

Zicheng Zhang, Tengchuan Kou, Shushi Wang, Chunyi Li, Wei Sun, Wei Wang, Xiaoyu Li, Zongyu Wang, Xuezhi Cao, Xiongkuo Min, Xiaohong Liu, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Meituan(美团)

Comments CVPR 2025 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20522 2025-06-17 cs.CV

MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks

Yifei Liu, Zhihang Zhong, Yifan Zhan, Sheng Xu, Xiao Sun

机构 * Shanghai AI Laboratory(上海人工智能实验室) Beihang University(北京航空航天大学) The University of Tokyo(东京大学)

Comments CVPR 2025; Project page:https://maskgaussian.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18688 2025-06-17 cs.CR cs.AI cs.LG

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment

Soumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan, Mengdi Wang, Alvaro Velasquez, Ahmad Beirami, Furong Huang, Dinesh Manocha, Amrit Singh Bedi

机构 * University of Maryland(马里兰大学) Indian Institute of Technology Bombay(印度班加罗尔理工学院) Princeton University(普林斯顿大学) University of Colorado Boulder(科罗拉多大学博尔德分校) Capital One(Capital One公司) University of Central Florida(佛罗里达中央大学)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00101 2025-06-17 cs.CV

SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail

Yingjun Du, Jiayi Shen, Xiantong Zhen, Cees G. M. Snoek

机构 * AIM Lab, University of Amsterdam(阿姆斯特丹大学AIM实验室) Inception Institute of Artificial Intelligence(Inception人工智能研究所) United Imaging Healthcare, Co., Ltd.(联合影象医疗科技有限公司)

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11543 2025-06-16 cs.CV cs.AI cs.LG

FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation

Zhuguanyu Wu, Shihe Wang, Jiayi Zhang, Jiaxin Chen, Yunhong Wang

机构 * State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, China(虚拟现实技术与系统国家重点实验室,北京航空航天大学) School of Computer Science and Engineering, Beihang University, Beijing, China(计算机科学与工程学院,北京航空航天大学)

Comments CVPR 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11417 2025-06-16 cs.CV cs.AI

Stop learning it all to mitigate visual hallucination, Focus on the hallucination target

Dokyoon Yoon, Youngsook Song, Woomyong Park

机构 * SIONIC AI

Comments Accepted to CVPR 2025

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11093 2025-06-16 cs.CV

EfficientQuant: An Efficient Post-Training Quantization for CNN-Transformer Hybrid Models on Edge Devices

Shaibal Saha, Lanyu Xu

机构 * Oakland University(奥克兰大学)

Comments Accepted to the 4th Workshop on Transformers for Vision (T4V) at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01881 2025-06-16 cs.CV cs.AI cs.LG cs.MM cs.RO

PhysNav-DG: A Novel Adaptive Framework for Robust VLM-Sensor Fusion in Navigation Applications

Trisanth Srinivasan, Santosh Patapati

机构 * Cyrion Labs(塞里昂实验室)

Comments Accepted at IEEE/CVF Computer Society Conference on Computer Vision and Pattern Recognition Workshops 2025 (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21099 2025-06-16 cs.CV

Learning Class Prototypes for Unified Sparse Supervised 3D Object Detection

Yun Zhu, Le Hui, Hang Yang, Jianjun Qian, Jin Xie, Jian Yang

机构 * PCA Lab, Nanjing University of Science and Technology(南京理工大学光电学院) School of Electronics and Information, Northwestern Polytechnical University(西北工业大学电子与信息学院) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学软件新技术国家重点实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05205 2025-06-16 cs.CV cs.AI

Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning

Xueyi Ke, Satoshi Tsutsui, Yayun Zhang, Bihan Wen

机构 * Nanyang Technological University(南洋理工大学) The Max Planck Institute for Psycholinguistics(马克斯·普朗克心理学语言学研究所)

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16318 2025-06-16 cs.CV cs.AI

One Diffusion to Generate Them All

Duong H. Le, Tuan Pham, Sangho Lee, Christopher Clark, Aniruddha Kembhavi, Stephan Mandt, Ranjay Krishna, Jiasen Lu

机构 * Allen Institute for AI(人工智能研究所) University of California, Irvine(加州大学尔湾分校) University of Washington(华盛顿大学)

Comments CVPR 2025; two first authors contribute equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16271 2025-06-16 cs.CV

FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned Priors

Chin-Yang Lin, Chung-Ho Wu, Chang-Han Yeh, Shih-Han Yen, Cheng Sun, Yu-Lun Liu

Comments Paper accepted to CVPR 2025. Project page: https://linjohnss.github.io/frugalnerf/

详情

展开后加载摘要…

URL PDF HTML 收藏