arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11876
2312.09181 2026-02-13 cs.CV

Improving Efficiency of Diffusion Models via Multi-Stage Framework and Tailored Multi-Decoder Architectures

通过多阶段框架和定制多解码器架构提高扩散模型的效率

Huijie Zhang, Yifu Lu, Ismail Alkhouri, Saiprasad Ravishankar, Dogyoon Song, Qing Qu

机构 * University of Michigan(密歇根大学) Michigan State University(密歇根州立大学)

AI总结 本文提出多阶段框架和定制多解码器架构,以提高扩散模型的训练和采样效率。

Comments The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11005 2026-02-12 cs.CV

Interpretable Vision Transformers in Monocular Depth Estimation via SVDA

通过SVDA实现单目深度估计中的可解释性视觉变换器

Vasileios Arampatzakis, George Pavlidis, Nikolaos Mitianoudis, Nikos Papamarkos

机构 * Dept. Electrical and Computer Engineering(电子与计算机工程系) Democritus University of Thrace(德米特里乌斯大学) Athena Research Center(雅典研究中心) University Campus at Kimmeria(基米里亚大学校园)

AI总结 SVDA通过引入谱结构化的注意力机制,提升了单目深度估计的可解释性,同时保持了预测精度并降低了计算开销。

Comments 8 pages, 2 figures, submitted to CVPR Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19844 2026-02-12 cs.CV

ProAPO: Progressively Automatic Prompt Optimization for Visual Classification

ProAPO:逐步自动提示优化用于视觉分类

Xiangyan Qu, Gaopeng Gou, Jiamin Zhuang, Jing Yu, Kun Song, Qihao Wang, Yili Li, Gang Xiong

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) School of Information Engineering, Minzu University of China(民族大学信息工程学院) University of Science and Technology Beijing(北京科技大学)

AI总结 ProAPO通过逐步自动优化提示,提升视觉分类性能,优于现有方法并在多个数据集上取得显著改进。

Comments Accepted to the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03507 2026-02-11 cs.CV

REFA: Real-time Egocentric Facial Animations for Virtual Reality

REFA: 虚拟现实中的实时第一人称面部动画

Qiang Zhang, Tong Xiao, Haroun Habeeb, Larissa Laich, Sofien Bouaziz, Patrick Snape, Wenjing Zhang, Matthew Cioffi, Peizhao Zhang, Pavel Pidlypenskyi, Winnie Lin, Luming Ma, Mengjiao Wang, Kunpeng Li, Chengjiang Long, Steven Song, Martin Prazak, Alexander Sjoholm, Ajinkya Deogade, Jaebong Lee, Julio Delgado Mangas, Amaury Aubel

机构 * Reality Labs at Meta(Meta 实验室)

AI总结 REFA通过基于蒸馏的方法实现虚拟现实中的实时面部表情跟踪,利用异构数据训练模型,提升虚拟角色表达的非侵入性与准确性。

Comments CVPR 2024 Workshop

Journal ref 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15592 2026-02-09 cs.CV

VP Lab: a PEFT-Enabled Visual Prompting Laboratory for Semantic Segmentation

VP Lab: 一种支持PEFT的视觉提示实验室用于语义分割

Niccolo Avogaro, Thomas Frick, Yagmur G. Cinar, Daniel Caraballo, Cezary Skura, Filip M. Janicki, Piotr Kluska, Brown Ebouky, Nicola Farronato, Florian Scheidegger, Cristiano Malossi, Konrad Schindler, Andrea Bartezzaghi, Roy Assaf, Mattia Rigotti

机构 * IBM Research(IBM研究院) ETH Zürich(苏黎世联邦理工学院)

AI总结 VP Lab通过E-PEFT技术提升视觉提示的语义分割性能,实现高效且交互式的模型部署。

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05598 2026-02-06 cs.CV cs.AI

CAViT -- Channel-Aware Vision Transformer for Dynamic Feature Fusion

CAViT -- 基于通道感知的视觉变换器用于动态特征融合

Aon Safdar, Mohamed Saadeldin

机构 * School of Computer Science, University College Dublin, Republic of Ireland(都柏林大学计算机科学学院)

AI总结 CAViT通过动态注意力机制提升视觉变换器的特征融合能力,实现更高的准确率和效率。

Comments Presented at the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025 (CVPR 25) in the 4th Workshop on Transformers for Visions - T4V (https://sites.google.com/view/t4v-cvpr25/) Accepted for Publication at 33rd International Conference on Artificial Intelligence and Cognitive Science (AICS 2025), where it was shortlisted for Best Paper Award. (https://aicsconf.org/?page_id=278)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09606 2026-02-03 cs.CV

Feat2GS: Probing Visual Foundation Models with Gaussian Splatting

Feat2GS: 通过高斯点云探测视觉基础模型

Yue Chen, Xingyu Chen, Anpei Chen, Gerard Pons-Moll, Yuliang Xiu

AI总结 Feat2GS通过高斯点云探测视觉基础模型的3D理解能力,无需3D数据即可分析几何和纹理意识。

Comments Project Page: https://fanegg.github.io/Feat2GS/

Journal ref Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20390 2026-02-03 cs.CV cs.GR cs.RO

InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions

InterMimic: 向物理基人类-物体交互的通用全身控制迈进

Sirui Xu, Hung Yu Ling, Yu-Xiong Wang, Liang-Yan Gui

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Electronic Arts(电子艺名)

AI总结 InterMimic通过课程策略和强化学习微调,实现从运动捕捉数据中学习并生成高质量的人类-物体交互。

Comments CVPR 2025. Project Page: https://sirui-xu.github.io/InterMimic/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00678 2026-02-02 cs.CV

2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image Classification

2DMamba:用于全像素全滑动图像分类的高效状态空间模型

Jingwei Zhang, Anh Tien Nguyen, Xi Han, Vincent Quoc-Huy Trinh, Hong Qin, Dimitris Samaras, Mahdi S. Hosseini

机构 * Stony Brook University(石溪大学) Korea University(韩国大学) Concordia University(康科迪亚大学) Mila–Quebec AI Institute(魁北克AI研究院) University of Montreal Hospital Center(蒙特利尔医院中心)

AI总结 2DMamba通过整合图像的2D空间结构,提出了一种高效的2D状态空间模型,用于全像素全滑动图像分类和自然图像处理,提升了多个指标的性能。

Comments Accepted in CVPR 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03623 2026-01-28 cs.CV

Bounding Box-Guided Diffusion for Synthesizing Industrial Images and Segmentation Map

边界框引导的扩散模型用于合成工业图像和分割图

Emanuele Caruso, Alessandro Simoni, Francesco Pelosin

机构 * Covision Lab(科维森实验室)

AI总结 本文提出基于扩散模型的边界框引导方法,用于生成高保真的工业图像和分割图,提升缺陷一致性和空间准确性,通过定量指标评估方法有效性,推动更可靠的分割模型发展。

Comments Accepted at Synthetic Data for Computer Vision Workshop - CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09665 2026-01-28 cs.CV

Revealing Subtle Phenotypes in Small Microscopy Datasets Using Latent Diffusion Models

利用潜在扩散模型揭示小样本显微图像中的细微表型

Anis Bourou, Biel Castaño Segade, Thomas Boyer, Valérie Mezger, Auguste Genovesio

机构 * Ecole Normale Supérieure(法国国家科学研究院) Université Paris Cité(巴黎cité大学)

AI总结 本文提出利用预训练的潜在扩散模型,在小样本显微图像中有效检测细微表型变化。

Comments Published to CVPR CVDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03167 2026-01-27 cs.CV cs.GR

CloSET: Modeling Clothed Humans on Continuous Surface with Explicit Template Decomposition

CloSET: 基于连续表面建模的着装人类建模

Hongwen Zhang, Siyou Lin, Ruizhi Shao, Yuxiang Zhang, Zerong Zheng, Han Huang, Yandong Guo, Yebin Liu

机构 * Tsinghua University(清华大学) OPPO Research Institute(OPPO研究院)

AI总结 CloSET通过分解显式服装模板并学习姿态依赖的褶皱,提升服装变形建模的精度和泛化能力。

Comments CVPR 2023 Paper, Update project page: https://zhanghongwen.cn/closet

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16788 2026-01-26 cs.CV cs.AI

REL-SF4PASS: Panoramic Semantic Segmentation with REL Depth Representation and Spherical Fusion

REL-SF4PASS:基于REL深度表示和球形融合的全景语义分割

Xuewei Li, Xinghan Bao, Zhimin Chen, Xi Li

机构 * School of Electronic and Information Engineering, Shanghai DianJi University(电子信息学院,上海电机大学) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

AI总结 REL-SF4PASS通过REL深度表示和球形动态多模态融合方法,提升全景语义分割的性能和鲁棒性。

Comments submitted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17372 2026-01-22 cs.CV

Coupled Laplacian Eigenmaps for Locally-Aware 3D Rigid Point Cloud Matching

耦合拉普拉斯特征映射用于局部感知的3D刚性点云匹配

Matteo Bastico, Etienne Decencière, Laurent Corté, Yannick Tillier, David Ryckelynck

机构 * Mines Paris, Université PSL(巴黎 Mines 学院,巴黎大学) Centre des Matériaux (MAT), UMR7633 CNRS(材料中心(MAT),CNRS UMR7633) Centre de Morphologie Mathématique (CMM)(数学形态学中心(CMM)) Centre de Mise en Forme des Matériaux (CEMEF), UMR7635 CNRS(材料成型中心(CEMEF),CNRS UMR7635)

AI总结 本文提出基于耦合拉普拉斯特征映射的3D点云匹配方法,通过考虑局部结构提升匹配精度,应用于物体异常检测和骨侧估计任务。

Comments This paper has been accepted at Computer Vision and Patter Recognition (CVPR) 2024

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 3447-3458

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13416 2026-01-21 cs.CV

Diffusion Representations for Fine-Grained Image Classification: A Marine Plankton Case Study

扩散表示用于细粒度图像分类:一个海洋浮游生物案例研究

A. Nieto Juscafresa, Á. Mazcuñán Herreros, J. Sullivan

机构 * KTH Royal Institute of Technology(皇家理工学院)

AI总结 本文提出利用冻结的扩散模型作为特征编码器,通过多层特征和时间步探测提升细粒度图像分类性能,在浮游生物监测中验证了其有效性。

Comments 21 pages, 6 figures, CVPR format

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13401 2026-01-21 cs.CV cs.AI

Reasoning with Pixel-level Precision: QVLM Architecture and SQuID Dataset for Quantitative Geospatial Analytics

基于像素级精度的推理:QVLM架构与SQuID数据集用于定量遥感分析

Peter A. Massih, Eric Cosatto

机构 * Department of Machine Learning, NEC Laboratories America(机器学习系,NEC美国实验室)

AI总结 本文提出QVLM架构和SQuID数据集,通过解耦语言理解和视觉分析,提升定量空间推理的准确性。

Comments Submitted to CVPR 2026. Introduces the QVLM architecture and the SQuID dataset for quantitative geospatial reasoning. Dataset DOI: 10.57967/hf/7565

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09823 2026-01-19 cs.CV

NanoSD: Edge Efficient Foundation Model for Real Time Image Restoration

NanoSD:面向边缘设备的高效图像修复扩散基础模型

Subhajit Sanyal, Srinivas Soumitri Miriyala, Akshay Janardan Bankar, Manjunath Arveti, Sowmya Vajrala, Shreyas Pandith, Sravanth Kodavanti, Abhishek Ameta, Harshit, Amit Satish Unde

机构 * Samsung Research India, Bangalore(三星印度研究, 波哥达)

AI总结 NanoSD通过网络手术、特征蒸馏和结构化扩展,构建了高效扩散基础模型家族,实现边缘设备上的实时图像修复与生成。

Comments Submitted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10228 2026-01-16 cs.CV cs.MM eess.IV

Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge

优化多模态大语言模型以实现视角视频理解:HD-EPIC VQA挑战的解决方案

Sicheng Yang, Yukai Huang, Shitong Sun, Weitong Cai, Jiankang Deng, Jifei Song, Zhensong Zhang

AI总结 本文提出一种优化多模态大语言模型的方法,通过预处理、微调和时间链式思考提示,在HD-EPIC VQA挑战中实现了41.6%的准确率。

Comments 4 pages, 1 figure, CVPR 2025 EgoVis Workshop, 2nd Place in HD-EPIC Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09406 2026-01-15 cs.CV cs.HC cs.LG cs.RO

Human-in-the-Loop Segmentation of Multi-species Coral Imagery

人机协同的多物种珊瑚图像分割

Scarlett Raine, Ross Marchant, Brano Kusy, Frederic Maire, Niko Suenderhauf, Tobias Fischer

机构 * QUT Centre for Robotics(QUT机器人中心) Image Analytics(图像分析) CSIRO Data61(CSIRO数据61)

AI总结 本文提出利用DINOv2特征和KNN实现高效珊瑚图像分割,仅用5个点标签即可在mIoU上提升13.5%。

Comments IEEE Journal of Oceanic Engineering accepted preprint of extended paper, 36 pages, 14 figures. Original conference paper (v2) accepted at the CVPR 2024 3rd Workshop on Learning with Limited Labelled Data for Image and Video Understanding (L3D-IVU)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08811 2026-01-14 cs.CV cs.AI

Reasoning Matters for 3D Visual Grounding

推理在3D视觉定位中至关重要

Hsiang-Wei Huang, Kuang-Ming Chen, Wenhao Chai, Cheng-Yen Yang, Jen-Hao Cheng, Jenq-Neng Hwang

机构 * University of Washington(华盛顿大学)

AI总结 本文提出了一种自动合成3D视觉定位数据的管道,并引入了在仅使用1.6%训练数据下表现优于现有方法的Reason3DVG-8B模型,证明了推理在3D视觉定位中的重要性。

Comments 2025 CVPR Workshop on 3D-LLM/VLA: Bridging Language, Vision and Action in 3D Environments

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07660 2026-01-13 cs.CV

StdGEN++: A Comprehensive System for Semantic-Decomposed 3D Character Generation

StdGEN++: 一种用于语义分解3D人物生成的综合性系统

Yuze He, Yanning Zhou, Wang Zhao, Jingwen Ye, Zhongkai Wu, Ran Yi, Yong-Jin Liu

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Tencent AIPD(腾讯AIPD) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)

AI总结 StdGEN++是一种基于双分支语义感知模型的综合性系统,通过语义分解和高效生成技术,实现高保真3D人物生成,适用于自动化角色资产生产。

Comments 13 pages, 12 figures. Extended version of CVPR 2025 paper arXiv:2411.05738

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07377 2026-01-13 cs.CV cs.AI

Learning Dynamic Collaborative Network for Semi-supervised 3D Vessel Segmentation

学习动态协作网络用于半监督3D血管分割

Jiao Xu, Xin Chen, Lihe Zhang

机构 * Dalian University of Technology(大连理工大学) City University of Hong Kong(香港城市大学)

AI总结 DiCo通过动态协作网络和多视图整合模块,提升半监督3D血管分割性能

Comments Accepted to the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17813 2026-01-13 cs.CV

CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair Loss

CLOC:基于多边距n对损失的有序分类对比学习

Dileepa Pitawela, Gustavo Carneiro, Hsiang-Ting Chen

机构 * University of Adelaide, Australia(阿德莱德大学) CVSSP, University of Surrey, UK(Surrey 大学计算机视觉与模式识别研究所)

AI总结 CLOC通过多边距n对损失提升有序分类性能,实现灵活决策边界和可解释的有序表示。

Comments Accepted in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01695 2026-01-06 cs.CV

Learnability-Driven Submodular Optimization for Active Roadside 3D Detection

基于可学习性的子模优化用于主动道路3D检测

Ruiyu Mao, Baoming Zhang, Nicholas Ruozzi, Yunhui Guo

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

AI总结 本研究提出了一种基于可学习性的主动学习框架,用于道路侧单目3D物体检测,通过选择信息丰富且可可靠标注的场景,有效减少标注成本并提升模型性能。

Comments 10 pages, 7 figures. Submitted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00991 2026-01-06 cs.CV

UnrealPose: Leveraging Game Engine Kinematics for Large-Scale Synthetic Human Pose Data

UnrealPose:利用游戏引擎动力学生成大规模合成人体姿态数据

Joshua Kawaguchi, Saad Manzur, Emily Gao Wang, Maitreyi Sinha, Bryan Vela, Yunxi Wang, Brandon Vela, Wayne B. Hayes

AI总结 UnrealPose通过Unreal Engine 5生成大规模合成人体姿态数据,包含100万帧,涵盖多样化的场景和动作,用于提升姿态估计和人体检测的性能。

Comments CVPR 2026 submission. Introduces UnrealPose-1M dataset and UnrealPose-Gen pipeline

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08745 2026-01-06 eess.IV cs.CV cs.MM

Enhancing Blind Video Quality Assessment with Rich Quality-aware Features

通过丰富的质量感知特征增强盲视频质量评估

Wei Sun, Linhan Cao, Jun Jia, Zhichao Zhang, Zicheng Zhang, Xiongkuo Min, Guangtao Zhai

机构 * East China Normal University(华东师范大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 本文提出RQ-VQA方法,通过整合多源特征提升盲视频质量评估的泛化能力与准确性。

Comments RQ-VQA won first place in the CVPR NTIRE 2024 Short-form UGC Video Quality Assessment Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22874 2025-12-30 cs.CV

Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of Samples

让样本说话:通过利用样本的聚类性来缓解虚假相关性

Weiwei Li, Junzhuo Liu, Yuanyuan Ren, Yuchen Zheng, Yahao Liu, Wen Li

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Shihezi University(石河子大学)

AI总结 本文提出了一种通过样本聚类性缓解深度学习模型中虚假相关性的方法,通过识别、中和、消除和更新四个步骤提升模型性能。

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20194 2025-12-24 cs.CV eess.IV

Generative Latent Coding for Ultra-Low Bitrate Image Compression

生成性潜在编码用于超低比特率图像压缩

Zhaoyang Jia, Jiahao Li, Bin Li, Houqiang Li, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

AI总结 本研究提出生成性潜在编码方法,通过在潜在空间中进行变换编码,实现超低比特率下的高真实感和高保真度图像压缩,并在多个数据集上验证了其有效性。

Comments Accepted at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20174 2025-12-24 cs.CV cs.CL cs.IR

Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark

迈向基于自然语言的文档图像检索:新数据集和基准

Hao Guo, Xugong Qin, Jun Jie Ou Yang, Peng Zhang, Gangyan Zeng, Yubo Li, Hailun Lin

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Science and Engineering, Nanjing University of Science and Technology(南京理工大学 cyber 科学与工程学院) State Key Laboratory of Cyberspace Security Defense(网络空间安全防御国家重点实验室) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) University of Southern California(美国南加州大学) Laboratory for Advanced Computing and Intelligence Engineering(先进计算与智能工程实验室)

AI总结 本文提出基于自然语言的文档图像检索基准,通过生成细粒度语义查询提升检索性能,推动视觉文档理解领域研究。

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16397 2025-12-19 cs.CV cs.AI cs.GR

Using Gaussian Splats to Create High-Fidelity Facial Geometry and Texture

利用高斯散点创建高质量的面部几何和纹理

Haodi He, Jihun Yu, Ronald Fedkiw

机构 * Epic Games, Stanford University(Epic Games、斯坦福大学)

AI总结 本文提出利用高斯散点重建高质量面部几何与纹理,通过约束三角化表面并解耦纹理与光照,实现高保真视觉效果。

Comments Submitted to CVPR 2026. 21 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏