arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11876
2512.16202 2025-12-19 cs.CV cs.AI

Open Ad-hoc Categorization with Contextualized Feature Learning

开放性即需分类与上下文化特征学习

Zilin Wang, Sangwoo Mo, Stella X. Yu, Sima Behpour, Liu Ren

机构 * University of Michigan(密歇根大学) UC Berkeley(加州大学伯克利分校) Bosch Center for AI(博世人工智能中心)

AI总结 OAK通过引入上下文标记和结合CLIP与GCD目标,实现了开放性即需分类的高准确率和可解释性。

Comments 26 pages, 17 figures

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18352 2025-12-19 cs.CV

Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models

Diffusion-4K: 基于潜在扩散模型的超高清图像合成

Jinjin Zhang, Qiuyu Huang, Junjie Liu, Xiefan Guo, Di Huang

机构 * State Key Laboratory of Complex and Critical Software Environment(复杂与关键软件环境国家重点实验室) Beihang University(北京航空航天大学) School of Computer Science and Engineering(计算机科学与工程学院)

AI总结 Diffusion-4K通过引入Aesthetic-4K基准和基于小波的微调方法,实现了高质量超高清图像合成。

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14711 2025-12-19 cs.CV

Low-Resolution Action Recognition for Tiny Actions Challenge

低分辨率动作识别挑战

Boyu Chen, Yu Qiao, Yali Wang

机构 * ShenZhen Key Lab of Computer Vision and Pattern Recognition, Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, China(深圳计算机视觉与模式识别重点实验室,深圳先进技术研究院,中国科学院,中国) University of Chinese Academy of Sciences(中国科学院大学) Shanghai AI Laboratory, Shanghai, China(上海人工智能实验室,上海,中国) SIAT Branch, Shenzhen Institute of Artificial Intelligence and Robotics for Society(SIAT分支,深圳人工智能与机器人社会研究院)

AI总结 本文提出了一种针对低分辨率动作识别挑战的解决方案,通过数据平衡、双分辨率蒸馏和模型集成提升长尾类别性能,取得排行榜第一。

Comments This article is the report of the CVPR 2022 ActivityNet workshop Tiny Actions Challenge(https://tinyactions-cvpr22.github.io/). The time of the first submission to the organizers is June 6th

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14542 2025-12-17 cs.CV

HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion

HiFi-Portrait: 零样本身份保留肖像生成的高保真多脸融合

Yifang Xu, Benxiang Zhai, Yunzhuo Sun, Ming Li, Yang Li, Sidan Du

机构 * Nanjing University(南京大学) Dalian University of Technology(大连理工大学) Nanjing University of Information Science and Technology(南京信息科技大学)

AI总结 HiFi-Portrait通过高保真多脸融合技术实现零样本身份保留肖像生成,提升了面部相似性和可控性,同时兼容SDXL工作。

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10031 2025-12-12 cs.CV cs.AI

ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image Objects

ABBSPO:自适应边界框缩放与对称先验的面向预测用于检测空中的图像对象

Woojin Lee, Hyugjae Chang, Jaeho Moon, Jaehyup Lee, Munchurl Kim

AI总结 ABBSPO通过自适应边界框缩放和对称先验角度损失,提升弱监督定向对象检测的准确性和效率。

Comments 17 pages, 11 figures, 8 tables, supplementary included. Accepted to CVPR 2025. Please visit our project page at https://kaist-viclab.github.io/ABBSPO_site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02912 2025-12-12 cs.CV cs.AI cs.GR cs.LG

ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware Prompts

ShapeWords: 基于3D形状感知提示的文本到图像合成

Dmitry Petrov, Pradyumn Goyal, Divyansh Shivashok, Yuanming Tao, Melinos Averkiou, Evangelos Kalogerakis

机构 * UMass Amherst(马萨诸塞大学阿默斯特分校) CYENS CoE(CYENS联合学院) University of Cyprus(塞浦路斯大学) TU Crete(希腊技术大学)

AI总结 ShapeWords通过融合3D形状信息与文本提示,提升文本到图像合成的准确性与3D结构感知能力。

Comments Project webpage: https://lodurality.github.io/shapewords/ (CVPR 2025 paper), this Author Accepted Manuscript version is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) in accordance with the ERC / Horizon Europe open-access mandate

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10378 2025-12-11 cs.CV cs.AI cs.CY cs.LG

Second Edition FRCSyn Challenge at CVPR 2024: Face Recognition Challenge in the Era of Synthetic Data

CVPR 2024 第二届合成数据时代人脸识别挑战赛:合成数据时代的人脸识别挑战

Ivan DeAndres-Tame, Ruben Tolosana, Pietro Melzi, Ruben Vera-Rodriguez, Minchul Kim, Christian Rathgeb, Xiaoming Liu, Aythami Morales, Julian Fierrez, Javier Ortega-Garcia, Zhizhou Zhong, Yuge Huang, Yuxi Mi, Shouhong Ding, Shuigeng Zhou, Shuai He, Lingzhi Fu, Heng Cong, Rongyu Zhang, Zhihong Xiao, Evgeny Smirnov, Anton Pimenov, Aleksei Grigorev, Denis Timoshenko, Kaleb Mesfin Asfaw, Cheng Yaw Low, Hao Liu, Chuyi Wang, Qing Zuo, Zhixiang He, Hatef Otroshi Shahreza, Anjith George, Alexander Unnervik, Parsa Rahimi, Sébastien Marcel, Pedro C. Neto, Marco Huber, Jan Niklas Kolf, Naser Damer, Fadi Boutros, Jaime S. Cardoso, Ana F. Sequeira, Andrea Atzori, Gianni Fenu, Mirko Marras, Vitomir Štruc, Jiang Yu, Zhangjie Li, Jichun Li, Weisong Zhao, Zhen Lei, Xiangyu Zhu, Xiao-Yu Zhang, Bernardo Biesseck, Pedro Vidal, Luiz Coelho, Roger Granada, David Menotti

AI总结 CVPR 2024第二届挑战赛探讨合成数据在人脸识别中的应用,旨在解决隐私、偏见和泛化能力等技术限制,通过新子任务推动面部生成方法的发展。

Comments arXiv admin note: text overlap with arXiv:2311.10476

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRw 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.05664 2025-12-11 cs.RO cs.AI

Altruistic Maneuver Planning for Cooperative Autonomous Vehicles Using Multi-agent Advantage Actor-Critic

为合作自主车辆的利他性动作规划使用多智能体优势Actor-Critic

Behrad Toghi, Rodolfo Valiente, Dorsa Sadigh, Ramtin Pedarsani, Yaser P. Fallah

AI总结 本文提出一种多智能体优势Actor-Critic算法,用于自动驾驶车辆在混合交通环境中的利他性动作规划,以提升交通效率与安全。

Comments Accepted to 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2021) - Workshop on Autonomous Driving: Perception, Prediction and Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.10496 2025-12-11 cs.CV cs.CL

Two Causal Principles for Improving Visual Dialog

为改进视觉对话的两个因果原则

Jiaxin Qi, Yulei Niu, Jianqiang Huang, Hanwang Zhang

机构 * Nanyang Technological University(南洋理工大学) Renmin University of China(中国人民大学) Damo Academy, Alibaba Group(阿里达摩院)

AI总结 本文提出两个因果原则以改进视觉对话模型,通过移除对话历史的直接输入和消除未观察到的混杂因素,提升模型性能。

Comments Accepted by CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23606 2025-12-09 cs.CV

Blurry-Edges: Photon-Limited Depth Estimation from Defocused Boundaries

模糊边缘:从模糊边界中进行光子限制深度估计

Wei Xu, Charles James Wagner, Junjie Luo, Qi Guo

机构 * Elmore Family School of Electrical and Computer Engineering(埃洛姆家庭电气与计算机工程学院)

AI总结 本文提出了一种基于Blurry-Edges表示的深度估计方法,通过深度神经网络从不同模糊程度的图像中预测深度信息,提升了光子受限图像的深度估计精度。

Comments Accepted to CVPR 2025. Project page: https://blurry-edges.qiguo.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13650 2025-12-09 eess.IV cs.CV

Tyche: Stochastic In-Context Learning for Medical Image Segmentation

Tyche: 医疗图像分割的随机上下文学习

Marianne Rakic, Hallee E. Wong, Jose Javier Gonzalez Ortiz, Beth Cimini, John Guttag, Adrian V. Dalca

机构 * CSAIL MIT(MIT计算机科学与人工智能实验室) MIT Broad Institute(MITBroad研究所) MGH(麻省总医院) MosaicML DataBricks MIT HMS, MGH(哈佛医学院、麻省总医院)

AI总结 Tyche通过上下文学习方法,无需重新训练即可为医疗图像分割生成随机预测,提升分割任务的灵活性和鲁棒性。

Comments Accepted at IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) 2024 as a highlight. Code available at https://github.com/mariannerakic/tyche

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02031 2025-12-08 cs.CV

Neural Eulerian Scene Flow Fields

神经欧拉场景流场

Kyle Vedder, Neehar Peri, Ishan Khatri, Siyi Li, Eric Eaton, Mehmet Kocamaz, Yue Wang, Zhiding Yu, Deva Ramanan, Joachim Pehserl

机构 * University of Pennsylvania(宾夕法尼亚大学) NVIDIA(英伟达) Carnegie Mellon University(卡内基梅隆大学)

AI总结 EulerFlow通过神经先验估计空间时间微分方程,实现高质量场景流估计,在多个领域表现优异,超越现有方法。

Comments Accepted to ICLR 2025. Winner of CVPR 2024 WoD Argoverse Scene Flow Challenge, Unsupervised Track. Project page at https://vedder.io/eulerflow

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04309 2025-12-05 cs.CV cs.CL

Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction

仅文本训练的图像描述生成:结合检索增强与模态差距校正

Rui Fonseca, Bruno Martins, Gil Rocha

机构 * INESC-ID, Instituto Superior Tecnico, University of Lisbon(INESC-ID,理工学院,里斯本大学)

AI总结 TOMCap通过检索增强和模态差距校正,实现无需对齐图像-文本对的纯文本训练图像描述生成。

Comments Submitted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01846 2025-12-03 cs.CV

UVGS: Reimagining Unstructured 3D Gaussian Splatting using UV Mapping

UVGS: 重新构想使用UV映射的无结构3D高斯点撒技术

Aashish Rai, Dilin Wang, Mihir Jain, Nikolaos Sarafianos, Kefan Chen, Srinath Sridhar, Aayush Prakash

机构 * Brown University(布朗大学) Meta Reality Labs(Meta现实实验室)

AI总结 本文提出UVGS,通过UV映射将无结构3D高斯点撒转化为结构化2D表示,利用现有2D模型高效建模3D数据,并实现可扩展的生成应用。

Comments https://ivl.cs.brown.edu/uvgs

Journal ref CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00532 2025-12-02 cs.CV cs.RO

Image Generation as a Visual Planner for Robotic Manipulation

图像生成作为机器人操作的视觉规划器

Ye Pang

机构 * Southern China University of Technology(华南理工大学)

AI总结 本文提出利用图像生成模型作为机器人视觉规划器,通过轻度微调实现时间连贯的机器人操作视频生成。

Comments 11 pages 9 figures Under review at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08303 2025-12-01 cs.CV

Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers

通过多模态视觉序列变压器推进语义未来预测

Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis

机构 * Archimedes, Athena Research Center(阿基米德研究中心) National Technical University of Athens(希腊国家技术大学) University of Crete(克里特大学) IACM-Forth(第四研究机构(IACM-Forth))

AI总结 FUTURIST通过多模态视觉序列变压器架构实现高效的多模态未来语义预测,提升预测精度并简化训练流程。

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11561 2025-12-01 cs.CV

Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution

利用大规模语言模型回归准确的图像质量评分使用分数分布

Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, Chao Dong

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Multimedia Laboratory, The Chinese University of Hong Kong(香港中文大学多媒体实验室) Shanghai AI Laboratory(上海人工智能实验室) Shenzhen University of Advanced Technology(深圳先进技术大学) CPII under InnoHK(创新香港下的CPII)

AI总结 本研究提出基于分布的DeQA-Score模型,通过离散化评分分布为软标签,提升图像质量评分的准确性和一致性。

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10979 2025-11-26 cs.CV cs.AI

VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?

VidComposition: MLLMs能否分析编译视频中的构成?

Yolo Y. Tang, Junjia Guo, Hang Hua, Susan Liang, Mingqian Feng, Xinyang Li, Rui Mao, Chao Huang, Jing Bi, Zeliang Zhang, Pooyan Fazli, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) Arizona State University(亚利桑那州立大学)

AI总结 VidComposition通过编排视频和电影级注释评估MLLMs对视频构成的理解能力,揭示了当前模型在复杂编译视频分析中的局限性。

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18570 2025-11-25 cs.CV cs.RO

PhysGS: Bayesian-Inferred Gaussian Splatting for Physical Property Estimation

PhysGS:基于贝叶斯推断的高斯点云法用于物理性质估计

Samarth Chopra, Jing Liang, Gershom Seneviratne, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Stanford University(斯坦福大学)

AI总结 PhysGS通过贝叶斯推断结合视觉线索和语言先验,实现对物理属性的密集估计,提升质量、硬度和摩擦误差的准确性。

Comments Submitted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07974 2025-11-25 cs.RO

Anomaly Detection in Autonomous Driving: A Survey

自动驾驶中的异常检测:综述

Daniel Bogdoll, Maximilian Nitsche, J. Marius Zöllner

机构 * FZI Research Center for Information Technology(FZI信息科技研究中心) KIT Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

AI总结 本文综述了自动驾驶中基于多种传感器数据的异常检测技术,系统化分类了检测方法和边缘案例级别,并指出现有研究的不足。

Comments Daniel Bogdoll and Maximilian Nitsche contributed equally. Accepted for publication at CVPR 2022 WAD workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16546 2025-11-21 cs.CV

Progressive Supernet Training for Efficient Visual Autoregressive Modeling

渐进式超网络训练用于高效的视觉自回归建模

Xiaoyue Chen, Yuling Shi, Kaiyuan Li, Huandong Wang, Yong Li, Xiaodong Gu, Xinlei Chen, Mingbao Lin

机构 * Tsinghua University, China(清华大学) Shanghai Jiao Tong University, China(上海交通大学) Rakuten, Singapore(Rakuten)

AI总结 VARiant通过渐进式超网络训练实现高效视觉自回归建模,减少内存消耗并提升部署灵活性。

Comments Submitted to CVPR 2025. 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10219 2025-11-21 cs.CV cs.GR cs.LG

PUP 3D-GS: Principled Uncertainty Pruning for 3D Gaussian Splatting

PUP 3D-GS: 3D高斯散射的原理性不确定性修剪

Alex Hanson, Allen Tu, Vasu Singla, Mayuka Jayawardhana, Matthias Zwicker, Tom Goldstein

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

AI总结 PUP 3D-GS通过原理性不确定性修剪技术,在更高压缩比下保持视觉质量和前景细节,提升渲染速度并优化图像质量。

Comments CVPR 2025, Project Page: https://pup3dgs.github.io/

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 5949-5958

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15308 2025-11-20 cs.CV

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Yilin Di, Joao F. Henriques, Daniel Cremers

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) Technical University of Munich(慕尼黑技术大学) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

Comments This paper builds upon and extends our earlier conference paper Text2Loc presented at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13135 2025-11-19 cs.CV

MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation

Junjie Yang, Yuhao Yan, Gang Wu, Yuxuan Wang, Ruoyu Liang, Xinjie Jiang, Xiang Wan, Fenglei Fan, Yongquan Zhang, Feiwei Qin, Changmiao Wang

机构 * South China University of Technology(华南理工大学) Sun Yat-sen University(中山大学) Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University of Finance & Economics(浙江财经大学) National University of Singapore(新加坡国立大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) City University of Hong Kong(香港城市大学)

Comments CVPR 2026 Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13001 2025-11-18 cs.CV

Medal S: Spatio-Textual Prompt Model for Medical Segmentation

Pengcheng Shi, Jiawei Chen, Jiaqi Liu, Xinglin Zhang, Tao Chen, Lei Li

机构 * Medical Image Insights, Shanghai, China(医学影像洞察,上海,中国) University of Washington, Seattle, WA, USA(华盛顿大学,西雅图,华盛顿州,美国) University of Waterloo, Waterloo, ON, Canada(滑铁卢大学,滑铁卢,安大略省,加拿大) Xi'an Jiaotong University, Xi'an, China(西安交通大学,西安,中国)

Comments Accepted by CVPR 2025 Workshop MedSegFM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07251 2025-11-18 cs.LG cs.AI cs.CR cs.CV

MOS-Attack: A Scalable Multi-objective Adversarial Attack Framework

Ping Guo, Cheng Gong, Xi Lin, Fei Liu, Zhichao Lu, Qingfu Zhang, Zhenkun Wang

机构 * City University of Hong Kong(香港城市大学) CityU Shenzhen Research Institute(城大深圳研究院) Southern University of Science and Technology(南方科技大学)

Comments Camera ready version of CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17663 2025-11-17 cs.CV

Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications

Junyi Ma, Xieyuanli Chen, Jiawei Huang, Jingyi Xu, Zhen Luo, Jintao Xu, Weihao Gu, Rui Ai, Hesheng Wang

机构 * Shanghai Jiao Tong University(上海交通大学) HAOMO.AI Technology Co., Ltd.(HAOMO.AI技术有限公司) National University of Defense Technology(国防科技大学) Beijing Institute of Technology(北京理工大学)

Comments Accepted to CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08207 2025-11-14 cs.CV cs.LG

DICE: Discrete Inversion Enabling Controllable Editing for Multinomial Diffusion and Masked Generative Models

Xiaoxiao He, Quan Dao, Ligong Han, Song Wen, Minhao Bai, Di Liu, Han Zhang, Martin Renqiang Min, Felix Juefei-Xu, Chaowei Tan, Bo Liu, Kang Li, Hongdong Li, Junzhou Huang, Faez Ahmed, Akash Srivastava, Dimitris Metaxas

机构 * Rutgers University(新泽西罗格斯大学) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) Red Hat AI Innovation(红帽AI创新) Google DeepMind(谷歌DeepMind) NYU(纽约大学) Walmart Global Tech(沃尔玛全球技术) NEC Labs America(NEC美国实验室) Massachusetts Institute of Technology(麻省理工学院) ANU(澳大利亚国立大学) UT Arlington(德克萨斯大学阿灵顿分校)

Comments Project webpage: https://hexiaoxiao-cs.github.io/DICE/. This paper was accepted to CVPR 2025 but later desk-rejected post camera-ready, due to a withdrawal from ICLR made 14 days before reviewer assignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01955 2025-11-12 cs.CV

Scene-Centric Unsupervised Panoptic Segmentation

Oliver Hahn, Christoph Reich, Nikita Araslanov, Daniel Cremers, Christian Rupprecht, Stefan Roth

机构 * TU Darmstadt(图宾根大学) TU Munich(慕尼黑工业大学) University of Oxford(牛津大学) MCML ELIZA

Comments To appear at CVPR 2025. Christoph Reich and Oliver Hahn - both authors contributed equally. Code: https://github.com/visinf/cups Project page: https://visinf.github.io/cups/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18108 2025-11-12 cs.CV

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

Jing Bi, Junjia Guo, Yunlong Tang, Lianggong Bruce Wen, Zhang Liu, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) Corning Inc(康宁公司)

Journal ref CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏