arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11876
2405.14343 2025-06-16 cs.CV

Efficient Visual State Space Model for Image Deblurring

Lingshun Kong, Jiangxin Dong, Jinhui Tang, Ming-Hsuan Yang, Jinshan Pan

机构 * Nanjing University of Science and Technology(南京理工大学) University of California, Merced(加州大学默塞德分校) Google(谷歌)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10790 2025-06-13 cs.CV

Human-Robot Navigation using Event-based Cameras and Reinforcement Learning

Ignacio Bugueno-Cordova, Javier Ruiz-del-Solar, Rodrigo Verschae

机构 * Department of Electrical Engineering, Universidad de Chile(智利大学电气工程系) Advanced Mining Technology Center (AMTC), Universidad de Chile(智利大学先进采矿技术中心) Institute of Engineering Sciences, Universidad de O’Higgins(奥希金斯大学工程科学研究所)

Comments https://ibugueno.github.io/hr-navigation-using-event-cameras-and-rl/

Journal ref 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); Fifth International Workshop on Event-Based Vision

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10328 2025-06-13 cs.CV cs.AI cs.LG

Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework

Sadia Kamal, Tim Oates, Joy Wan

机构 * Department of Computer Science, University of Maryland, Baltimore County(计算机科学系,马里兰大学巴尔的摩县分校) Department of Dermatology, Johns Hopkins University School of Medicine(皮肤科系,约翰霍普金斯大学医学院)

Comments Accepted at IEEE/CVF Computer Society Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10286 2025-06-13 cs.CV

HalLoc: Token-level Localization of Hallucinations for Vision Language Models

Eunkyu Park, Minyeong Kim, Gunhee Kim

机构 * Seoul National University(首尔国立大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10242 2025-06-13 cs.CV

DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos

Rajeev Yasarla, Shizhong Han, Hong Cai, Fatih Porikli

机构 * Qualcomm AI Research(高通人工智能研究)

Comments CVPR 2025 Workshop on Autonomous Driving

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10182 2025-06-13 cs.CV

Improving Personalized Search with Regularized Low-Rank Parameter Updates

Fiona Ryan, Josef Sivic, Fabian Caba Heilbron, Judy Hoffman, James M. Rehg, Bryan Russell

机构 * Georgia Tech(佐治亚理工学院) Adobe Research(Adobe研究) CIIRC CTU(查理大学CIIRC) UIUC(伊利诺伊大学厄巴纳-香槟分校)

Comments CVPR 2025 Highlight. Code: http://github.com/adobe-research/polar-vl

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04715 2025-06-13 cs.CV

Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model

Zelu Qi, Ping Shi, Chaoyang Zhang, Shuqi Wang, Fei Zhao, Da Pan, Zefeng Ying

机构 * Communication University of China(中国通信大学)

Comments This paper has been accepted by CVPR Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06470 2025-06-13 cs.CV cs.DC

Federated Unsupervised Visual Representation Learning via Exploiting General Content and Personal Style

Yuewei Yang, Jingwei Sun, Ang Li, Hai Li, Yiran Chen

机构 * Duke University(杜克大学)

Comments Reformat to CVPR format

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09989 2025-06-12 cs.CV

Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes

Yiming Dou, Wonseok Oh, Yuqing Luo, Antonio Loquercio, Andrew Owens

机构 * University of Michigan(密歇根大学) University of Pennsylvania(宾夕法尼亚大学)

Comments CVPR 2025, Project page: https://www.yimingdou.com/hearing_hands/ , Code: https://github.com/Dou-Yiming/hearing_hands/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09952 2025-06-12 cs.CV cs.AI

UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting

Ziyi Wang, Yanran Zhang, Jie Zhou, Jiwen Lu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02550 2025-06-12 cs.CV cs.AI

Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025

Qiaohui Chu, Haoyu Zhang, Yisen Feng, Meng Liu, Weili Guan, Yaowei Wang, Liqiang Nie

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Pengcheng Laboratory(鹏城实验室) Shandong Jianzhu University(山东建筑大学)

Comments The champion solution for the Ego4D Long-Term Action Anticipation Challenge at the CVPR EgoVis Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04459 2025-06-12 cs.CV

Question-Aware Gaussian Experts for Audio-Visual Question Answering

Hongyeob Kim, Inyoung Jung, Dayoon Suh, Youjia Zhang, Sangmin Lee, Sungeun Hong

机构 * Sungkyunkwan University(成均馆大学) Purdue University(普渡大学)

Comments CVPR 2025. Code is available at https://github.com/AIM-SKKU/QA-TIGER

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09473 2025-06-12 cs.CV

Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning

Cheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao, Bolin Ding, Jia Li

机构 * State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) Tongyi Lab, Alibaba Group(阿里云实验室)

Comments 10 pages, 6 figures, CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09411 2025-06-12 cs.CV cs.AI

Synthetic Human Action Video Data Generation with Pose Transfer

Vaclav Knapp, Matyas Bohacek

机构 * SSPS(社会学系) Stanford University(斯坦福大学)

Journal ref Synthetic Data for Computer Vision Workshop @ CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09369 2025-06-12 cs.CV

ScaleLSD: Scalable Deep Line Segment Detection Streamlined

Zeran Ke, Bin Tan, Xianwei Zheng, Yujun Shen, Tianfu Wu, Nan Xue

机构 * Wuhan University(武汉大学) Ant Group(蚂蚁集团) NC State University(北卡罗来纳州立大学)

Comments accepted to CVPR 2025; 17 pages, appendices included

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09343 2025-06-12 cs.CV cs.RO

CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation

Yuxing Long, Jiyao Zhang, Mingjie Pan, Tianshu Wu, Taewhan Kim, Hao Dong

机构 * CFCS, School of Computer Science, Peking University(计算机学院,北京大学) PKU-Agibot Lab(北京大学Agibot实验室)

Comments CVPR 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09075 2025-06-12 cs.GR cs.CV cs.LG

SILK: Smooth InterpoLation frameworK for motion in-betweening A Simplified Computational Approach

Elly Akhoundi, Hung Yu Ling, Anup Anand Deshmukh, Judith Butepage

机构 * SEED, Electronic Arts(SEED电子艺术公司) Electronic Arts(电子艺术公司)

Comments Accepted to CVPR 2025 Human Motion Generation Workshop. 10 pages, 3 figures, 5 Tables, and 40 References

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09068 2025-06-12 cs.CV cs.LG cs.RO

BG-HOP: A Bimanual Generative Hand-Object Prior

Sriram Krishna, Sravan Chittupalli, Sungjae Park

机构 * Carnegie Mellon University(卡内基梅隆大学)

Comments Presented at Agents in Interaction, from Humans to Robots, CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08279 2025-06-12 cs.CV

SmartEraser: Remove Anything from Images using Masked-Region Guidance

Longtao Jiang, Zhendong Wang, Jianmin Bao, Wengang Zhou, Dongdong Chen, Lei Shi, Dong Chen, Houqiang Li

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

Comments Project at: https://longtaojiang.github.io/smarteraser.github.io/

Journal ref The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04540 2025-06-12 cs.LG cs.AI cs.CV cs.MA cs.RO

Sim-to-Real Causal Transfer: A Metric Learning Approach to Causally-Aware Interaction Representations

Ahmad Rahimi, Po-Chien Luan, Yuejiang Liu, Frano Rajič, Alexandre Alahi

机构 * École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院) Stanford University(斯坦福大学) ETH Zurich(苏黎世联邦理工学院)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08964 2025-06-11 cs.CV

ORIDa: Object-centric Real-world Image Composition Dataset

Jinwoo Kim, Sangmin Han, Jinho Jeong, Jiwoo Choi, Dongyoung Kim, Seon Joo Kim

机构 * Yonsei University(延世大学)

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08887 2025-06-11 cs.CV

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval

Leqi Shen, Guoqiang Gong, Tianxiang Hao, Tao He, Yifeng Zhang, Pengzhang Liu, Sicheng Zhao, Jungong Han, Guiguang Ding

机构 * School of Software(软件学院) BNRist Department of Automation, Tsinghua University(自动化系,清华大学) Hangzhou Zhuoxi Institute of Brain and Intelligence(杭州卓溪脑科学与智能研究院) GRG Banking Equipment Co., Ltd.(GRG银行设备有限公司) South China University of Technology(华南理工大学)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08189 2025-06-11 cs.CV cs.CL

Open World Scene Graph Generation using Vision Language Models

Amartya Dutta, Kazi Sajeed Mehrab, Medha Sawhney, Abhilash Neog, Mridul Khurana, Sepideh Fatemi, Aanish Pradhan, M. Maruf, Ismini Lourentzou, Arka Daw, Anuj Karpatne

Comments Accepted in CVPR 2025 Workshop (CVinW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15256 2025-06-11 cs.CV cs.AI

Zero-Shot Gaze-based Volumetric Medical Image Segmentation

Tatyana Shmykova, Leila Khaertdinova, Ilya Pershin

机构 * Research Center of the Artificial Intelligence Institute(人工智能研究所研究中心)

Comments Accepted to MMFM-BIOMED Workshop @ CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00788 2025-06-11 cs.CV

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

Wufei Ma, Luoxin Ye, Celso M de Melo, Jieneng Chen, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

Comments CVPR 2025 highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03896 2025-06-11 cs.CV

Multimodal Rationales for Explainable Visual Question Answering

Kun Li, George Vosselman, Michael Ying Yang

机构 * University of Twente(特文特大学) University of Bath(巴斯大学)

Comments Accepted to CVPR workshops 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08023 2025-06-11 q-bio.BM cs.AI cs.CE cs.CV cs.LG

Aligning Proteins and Language: A Foundation Model for Protein Retrieval

Qifeng Wu, Zhengzhe Liu, Han Zhu, Yizhou Zhao, Daisuke Kihara, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Purdue University(普渡大学)

Comments 4 pages for body, 3 pages for appendix, 11 figures. Accepted to CVPR 2025 Workshop on Multimodal Foundation Models for Biomedicine: Challenges and Opportunities(MMFM-BIOMED)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19009 2025-06-11 cs.CV cs.IR

Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval

Arun Reddy, Alexander Martin, Eugene Yang, Andrew Yates, Kate Sanders, Kenton Murray, Reno Kriz, Celso M. de Melo, Benjamin Van Durme, Rama Chellappa

机构 * Johns Hopkins Applied Physics Laboratory(约翰霍普金斯应用物理实验室) Johns Hopkins University(约翰霍普金斯大学) Human Language Technology Center of Excellence(人类语言技术卓越中心) DEVCOM Army Research Laboratory(陆军研究实验室)

Comments Accepted at CVPR 2025. 13 pages, 4 figures. Approved for public release: distribution unlimited

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07306 2025-06-11 cs.CV cs.AI cs.CL cs.LG cs.RO

TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation

Navid Rajabi, Jana Kosecka

机构 * George Mason University(乔治·马歇尔大学)

Comments Accepted to CVPR 2025 Workshop - Foundation Models Meet Embodied Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10494 2025-06-11 cs.CV cs.AI cs.LG cs.PF

SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

Yushu Wu, Zhixing Zhang, Yanyu Li, Yanwu Xu, Anil Kag, Yang Sui, Huseyin Coskun, Ke Ma, Aleksei Lebedev, Ju Hu, Dimitris Metaxas, Yanzhi Wang, Sergey Tulyakov, Jian Ren

机构 * Snap Inc. Northeastern University(东北大学) Rutgers University(罗格斯大学)

Comments https://snap-research.github.io/snapgen-v/

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏