arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11877
2504.01019 2025-04-02 cs.CV

MixerMDM: Learnable Composition of Human Motion Diffusion Models

Pablo Ruiz-Ponce, German Barquero, Cristina Palmero, Sergio Escalera, José García-Rodríguez

Comments CVPR 2025 Accepted - Project Page: https://pabloruizponce.com/papers/MixerMDM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14519 2025-04-02 cs.RO

Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning

Jiange Yang, Haoyi Zhu, Yating Wang, Gangshan Wu, Tong He, Limin Wang

Comments Accepted to CVPR 2025. Code Page: https://github.com/MCG-NJU/Tra-MoE

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00773 2025-04-02 cs.CV

DropGaussian: Structural Regularization for Sparse-view Gaussian Splatting

Hyunwoo Park, Gun Ryu, Wonjun Kim

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00665 2025-04-02 cs.CV

Monocular and Generalizable Gaussian Talking Head Animation

Shengjie Gong, Haojie Li, Jiapeng Tang, Dongming Hu, Shuangping Huang, Hao Chen, Tianshui Chen, Zhuoman Liu

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00660 2025-04-02 cs.LG

Learning to Normalize on the SPD Manifold under Bures-Wasserstein Geometry

Rui Wang, Shaocheng Jin, Ziheng Chen, Xiaoqing Luo, Xiao-Jun Wu

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00557 2025-04-02 cs.CV cs.LG

Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features

Jewon Lee, Ki-Ung Song, Seungmin Yang, Donguk Lim, Jaeyeon Kim, Wooksu Shin, Bo-Kyeong Kim, Yong Jae Lee, Tae-Ho Kim

Comments accepted at CVPR 2025 Workshop on ELVM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00527 2025-04-02 cs.CV

SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning

Fida Mohammad Thoker, Letian Jiang, Chen Zhao, Bernard Ghanem

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00430 2025-04-02 cs.CV

Data Synthesis with Diverse Styles for Face Recognition via 3DMM-Guided Diffusion

Yuxi Mi, Zhizhou Zhong, Yuge Huang, Qiuyang Yuan, Xuan Zhao, Jianqing Xu, Shouhong Ding, ShaoMing Wang, Rizen Guo, Shuigeng Zhou

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00380 2025-04-02 cs.CV

Hierarchical Flow Diffusion for Efficient Frame Interpolation

Yang Hai, Guo Wang, Tan Su, Wenjie Jiang, Yinlin Hu

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00379 2025-04-02 cs.CV

MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving

Zhiyuan Zhang, Xiaofan Li, Zhihao Xu, Wenjie Peng, Zijian Zhou, Miaojing Shi, Shuangping Huang

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00247 2025-04-02 cs.CV cs.AI

MultiMorph: On-demand Atlas Construction

S. Mazdak Abulnaga, Andrew Hoopes, Neel Dey, Malte Hoffmann, Marianne Rakic, Bruce Fischl, John Guttag, Adrian Dalca

Comments accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00219 2025-04-02 cs.CV

LITA-GS: Illumination-Agnostic Novel View Synthesis via Reference-Free 3D Gaussian Splatting and Physical Priors

Han Zhou, Wei Dong, Jun Chen

Comments Accepted by CVPR 2025. 3DGS, Adverse illumination conditions, Reference-free, Physical priors

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00185 2025-04-02 cs.CV cs.LG

Self-Evolving Visual Concept Library using Vision-Language Critics

Atharva Sehgal, Patrick Yuan, Ziniu Hu, Yisong Yue, Jennifer J. Sun, Swarat Chaudhuri

Comments CVPR camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00072 2025-04-02 cs.CV

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

Lucas Ventura, Antoine Yang, Cordelia Schmid, Gül Varol

Comments CVPR 2025 Camera ready. Project page: https://imagine.enpc.fr/~lucas.ventura/chapter-llama/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22420 2025-04-02 cs.CV

Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis

Jiangyong Huang, Baoxiong Jia, Yan Wang, Ziyu Zhu, Xiongkun Linghu, Qing Li, Song-Chun Zhu, Siyuan Huang

Comments CVPR 2025. Project page: https://beacon-3d.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21442 2025-04-02 cs.GR cs.CV

RainyGS: Efficient Rain Synthesis with Physically-Based Gaussian Splatting

Qiyu Dai, Xingyu Ni, Qianfan Shen, Wenzheng Chen, Baoquan Chen, Mengyu Chu

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17940 2025-04-02 cs.CV

FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation

Dong Zhao, Jinlong Li, Shuang Wang, Mengyao Wu, Qi Zang, Nicu Sebe, Zhun Zhong

Journal ref Conference on Computer Vision and Pattern Recognition 2025 Conference on Computer Vision and Pattern Recognition 2025 Conference on Computer Vision and Pattern Recognition 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10437 2025-04-02 cs.CV

4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models

Wanhua Li, Renping Zhou, Jiawei Zhou, Yingwei Song, Johannes Herter, Minghan Qin, Gao Huang, Hanspeter Pfister

Comments CVPR 2025. Project Page: https://4d-langsplat.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18098 2025-04-02 cs.CV cs.LG

Disentangling Safe and Unsafe Corruptions via Anisotropy and Locality

Ramchandran Muthukumar, Ambar Pal, Jeremias Sulam, Rene Vidal

Comments Published at IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025. Updated Acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19867 2025-04-02 cs.CV cs.AI cs.LG

Data-Free Group-Wise Fully Quantized Winograd Convolution via Learnable Scales

Shuokai Pan, Gerti Tuzi, Sudarshan Sreeram, Dibakar Gope

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16855 2025-04-02 cs.CL cs.IR

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Xin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li, Ziqi Dai, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li, Min Zhang

Comments Accepted to CVPR 2025, models at https://huggingface.co/Alibaba-NLP/gme-Qwen2-VL-2B-Instruct

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03735 2025-04-02 cs.CV

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Chaoyu Li, Eun Woo Im, Pooyan Fazli

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01095 2025-04-02 cs.AI cs.CV cs.LG

VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models

Muchao Ye, Weiyang Liu, Pan He

Comments Accepted in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00596 2025-04-02 cs.CV cs.AI

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation

Qiyao Xue, Xiangyu Yin, Boyuan Yang, Wei Gao

Comments 28 pages

Journal ref in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16801 2025-04-02 cs.CV

Controllable Human Image Generation with Personalized Multi-Garments

Yisol Choi, Sangkyung Kwak, Sihyun Yu, Hyungwon Choi, Jinwoo Shin

Comments CVPR 2025. Project page: https://omnious.github.io/BootComp

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24382 2025-04-01 cs.CV

Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Views

Chong Bao, Xiyu Zhang, Zehao Yu, Jiale Shi, Guofeng Zhang, Songyou Peng, Zhaopeng Cui

Comments Accepted to CVPR 2025. Project Page: https://zju3dv.github.io/free360/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24374 2025-04-01 cs.CV

ERUPT: Efficient Rendering with Unposed Patch Transformer

Maxim V. Shugaev, Vincent Chen, Maxim Karrenbach, Kyle Ashley, Bridget Kennedy, Naresh P. Cuntoor

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24210 2025-04-01 cs.CV cs.AI cs.MM

DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting

Seungjun Lee, Gim Hee Lee

Comments CVPR 2025. Project Page: https://diet-gs.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23882 2025-04-01 cs.CV

GLane3D : Detecting Lanes with Graph of 3D Keypoints

Halil İbrahim Öztürk, Muhammet Esat Kalfaoğlu, Ozsel Kilinc

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20308 2025-04-01 cs.GR cs.CV

Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metrics

Lee Chae-Yeon, Oh Hyun-Bin, Han EunGi, Kim Sung-Bin, Suekyeong Nam, Tae-Hyun Oh

Comments CVPR 2025. Project page: https://perceptual-3d-talking-head.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏