arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

共收录 11876
2509.21029 2026-03-03 cs.LG

FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction

FORCE:通过特征过度依赖校正实现可转移的视觉劫持攻击

Runqi Lin, Alasdair Paren, Suqin Yuan, Muyang Li, Philip Torr, Adel Bibi, Tongliang Liu

机构 * Sydney AI Centre, The University of Sydney(悉尼人工智能中心,悉尼大学) Department of Engineering Science, University of Oxford(工程科学系,牛津大学)

AI总结 FORCE方法通过校正特征过度依赖,提升视觉劫持攻击的跨模型可转移性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06719 2026-03-03 cs.CV cs.AI

Improving Wildlife Out-of-Distribution Detection: Africas Big Five

提升野生动物分布外检测:非洲五大动物

Mufhumudzi Muthivhi, Jiahao Huo, Fredrik Gustafsson, Terence L. van Zyl

机构 * Institute for Artificial Intelligent Systems(人工智能系统研究所) University of Johannesburg(约翰内斯堡大学) Department of Electrical Engineering(电气工程系) Linköping University(利马大学)

AI总结 本研究通过改进的NCM和对比学习方法,提升对非洲五大动物的分布外检测性能,实现AUPR-IN、AUPR-OUT和AUTC指标的显著提升。

Comments Presented at the CV4Animals Workshop at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

Journal ref CVPR 2025 Workshop on Computer Vision for Animal Behavior Tracking and Modeling (CV4Animals)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17132 2026-03-03 cs.CV cs.CL

Dynamic Token Reweighting for Robust Vision-Language Models

动态令牌重加权用于鲁棒的视觉-语言模型

Tanqiu Jiang, Jiacheng Liang, Rongyi Zhu, Jiawei Zhou, Fenglong Ma, Ting Wang

机构 * Stony Brook University(石溪大学) Pennsylvania State University(宾夕法尼亚州立大学)

AI总结 DTR通过优化KV缓存动态调整视觉令牌权重,提升视觉-语言模型对多模态劫持攻击的鲁棒性,同时保持模型性能。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02825 2026-03-03 cs.CV

Towards Application-Specific Evaluation of Vision Models: Case Studies in Ecology and Biology

面向应用的视觉模型评估:生态与生物学中的案例研究

Alex Hoi Hang Chan, Otto Brookes, Urs Waldmann, Hemal Naik, Iain D. Couzin, Majid Mirmehdi, Noël Adiko Houa, Emmanuelle Normand, Christophe Boesch, Lukas Boesch, Mimi Arandjelovic, Hjalmar Kühl, Tilo Burghardt, Fumihiro Kano

机构 * Centre for the Advanced Study of Collective Behaviour, University of Konstanz(康斯坦茨大学集体行为高级研究所以及大学) Department of Collective Behavior, Max Planck Institute of Animal Behavior(动物行为马克斯·普朗克研究所集体行为系) Department of Biology, University of Konstanz(康斯坦茨大学生物学系) School of Computer Science, University of Bristol(布里斯托大学计算机科学学院) Wild Chimpanzee Foundation(野生 chimpanzee 基金会) Department of Computer and Information Science, University of Konstanz(康斯坦茨大学计算机与信息科学系) Department of Ecology of Animal Societies, Max Planck Institute of Animal Behavior(动物行为马克斯·普朗克研究所动物社会生态学系) Senckenberg Museum of Natural History Goerlitz(戈尔茨自然史塞肯贝格博物馆) Max Planck Institute for Evolutionary Anthropology(进化人类学马克斯·普朗克研究所)

AI总结 本文提出通过应用特定指标评估视觉模型在生态与生物学中的实际效果,通过两个案例研究展示模型性能对下游分析的影响。

Comments Accepted at CVPR Workshops, CV4Animals 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.06373 2026-03-03 cs.CV

Towards Accurate One-Stage Object Detection with AP-Loss

通过AP损失实现更准确的一阶段目标检测

Kean Chen, Jianguo Li, Weiyao Lin, John See, Ji Wang, Lingyu Duan, Zhibo Chen, Changwei He, Junni Zou

AI总结 本文提出通过AP损失优化一阶段目标检测器,解决分类损失的不平衡问题,提升检测性能。

Comments 13 pages, 7 figures, 4 tables, main paper + supplementary material, accepted to CVPR 2019

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00207 2026-03-03 cs.CV cs.AI

VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models

VisRef: 通过思考进行视觉再聚焦以提高多模态大推理模型的测试时间扩展

Soumya Suvra Ghosal, Youngeun Kim, Zhuowei Li, Ritwick Chaudhry, Linghan Xu, Hongjing Zhang, Jakub Zablocki, Yifan Xing, Qin Zhang

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Amazon(亚马逊) Physion Labs(Physion实验室)

AI总结 VisRef通过视觉再聚焦提升多模态大推理模型测试时间扩展性能,有效解决视觉依赖任务中推理性能下降问题。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00149 2026-03-03 cs.CV cs.AI

Physics-Consistent Diffusion for Efficient Fluid Super-Resolution via Multiscale Residual Correction

基于多尺度残差校正的物理一致扩散用于高效流体超分辨率

Zhihao Li, Shengwei Dong, Chuang Yi, Junxuan Gao, Zhilu Lai, Zhiqiang Liu, Wei Wang, Guangtao Zhang

AI总结 ReMD通过多网格残差校正和多小波多尺度建模实现高效流体超分辨率,提升精度和频谱保真度,减少发散并降低采样步骤。

Comments Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24233 2026-03-02 cs.CV

Enhancing Spatial Understanding in Image Generation via Reward Modeling

通过奖励建模增强图像生成中的空间理解

Zhenyu Tang, Chaoran Feng, Yufan Deng, Jie Wu, Xiaojie Li, Rui Wang, Yunpeng Chen, Daquan Zhou

机构 * Peking University(北京大学) ByteDance Seed(字节跳动种子)

AI总结 本文提出SpatialScore奖励模型,通过增强空间理解能力提升图像生成的质量和效果。

Comments Accepted at CVPR 2026. Github: https://github.com/DAGroup-PKU/SpatialT2I Project website: https://dagroup-pku.github.io/SpatialT2I/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24222 2026-03-02 cs.CV cs.LG

MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy

MuViT:多分辨率视觉变换器用于跨尺度学习的显微镜分析

Albert Dominguez Mantes, Gioele La Manno, Martin Weigert

机构 * Swiss Federal Institute of Technology Lausanne(瑞士联邦理工学院洛桑分校) Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI)(可扩展数据与人工智能中心(ScaDS.AI)) Technische Universität Dresden(德累斯顿技术大学)

AI总结 MuViT 通过多分辨率视觉变换器在显微镜分析中实现跨尺度学习,提升多分辨率信息利用能力。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24148 2026-03-02 cs.CV

HumanOrbit: 3D Human Reconstruction as 360° Orbit Generation

HumanOrbit:3D人体重建作为360°轨道生成

Keito Suzuki, Kunyao Chen, Lei Wang, Bang Du, Runfa Blark Li, Peng Liu, Ning Bi, Truong Nguyen

机构 * University of California, San Diego(加州大学圣地亚哥分校) Qualcomm(高通公司)

AI总结 HumanOrbit通过视频扩散模型实现多视角人体生成,生成360°轨道视频并重建高保真3D模型。

Comments CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24059 2026-03-02 cs.CV cs.AI

Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization

量化专家:基于专家混合的令牌感知自适应误差重建用于大视觉-语言模型量化

Chenwei Jia, Baoting Li, Xuchong Zhang, Mingzhuo Wei, Bochen Lin, Hongbin Sun

机构 * State Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) Xi’an Jiaotong University(西安交通大学)

AI总结 本文提出Quant Experts,通过令牌感知的专家混合方法,提升大视觉-语言模型在量化过程中的任务准确性,保持与全精度模型相当的性能。

Comments 13 pages, 6 figures, including appendix, Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24020 2026-03-02 cs.CV

SR3R: Rethinking Super-Resolution 3D Reconstruction With Feed-Forward Gaussian Splatting

SR3R: 重新思考超分辨率3D重建与前馈高斯点散射

Xiang Feng, Xiangbo Wang, Tieshi Zhong, Chengkai Wang, Yiting Zhao, Tianxiang Xu, Zhenzhong Kuang, Feiwei Qin, Xuefei Yin, Yanming Zhu

机构 * Hangzhou Dianzi University(杭州电子科技大学) ShanghaiTech University(上海科技大学) Griffith University(格里菲斯大学) Peking University(北京大学)

AI总结 SR3R通过前馈高斯点散射框架,直接从稀疏低分辨率视角生成高分辨率3D重建,提升跨场景泛化能力和实时性能。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24014 2026-03-02 cs.CV cs.AI

Interpretable Debiasing of Vision-Language Models for Social Fairness

可解释的视觉-语言模型社会公平性偏见消除

Na Min An, Yoonna Jang, Yusuke Hirota, Ryo Hachiuma, Isabelle Augenstein, Hyunjung Shim

机构 * KAIST AI(韩国科学技术院人工智能研究所) University of Copenhagen(哥本哈根大学) NVIDIA(英伟达公司)

AI总结 本研究提出DeBiasLens框架,通过稀疏自编码器局部化VLM中的社会属性神经元,以可解释的方式消除模型中的社会偏见,同时保持语义知识。

Comments 25 pages, 30 figures, 13 Tables Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23980 2026-03-02 cs.CV

Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping

Venus: 多模态大语言模型在审美指导与裁剪中的基准测试与赋能

Tianxiang Du, Hulingxiao He, Yuxin Peng

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究所)

AI总结 Venus通过两阶段框架提升多模态大语言模型的审美指导与裁剪能力,实现可解释的审美优化。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23814 2026-03-02 cs.CV

Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation

基于3D几何先验的双臂操作动作预测

Chongyang Xu, Haipeng Li, Shen Cheng, Jingyu Hu, Haoqiang Fan, Ziliang Feng, Shuaicheng Liu

机构 * College of Computer Science, Sichuan University, China(四川大学计算机学院) University of Electronic Science and Technology of China, China(电子科技大学) Dexmal The Chinese University of Hong Kong, HK SAR, China(香港中文大学(HK SAR))

AI总结 本文提出基于3D几何先验的双臂操作预测框架,通过融合几何特征与2D语义信息,实现高效动作预测与空间理解。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23739 2026-03-02 cs.CV

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

U-Mind:一种实时光学交互的统一框架

Xiang Deng, Feng Gao, Yong Zhang, Youxin Pang, Xu Xiaoming, Zhuoliang Kang, Xiaoming Wei, Yebin Liu

机构 * Tsinghua University(清华大学) Meituan(美团)

AI总结 U-Mind 提出了一种统一框架,通过实时生成和多模态联合建模,实现高效且连贯的多模态交互,推动智能对话代理的发展。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23734 2026-03-02 cs.CV cs.CL

UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking

UTPTrack: 向简单和统一的令牌剪枝迈进:视觉跟踪中的统一令牌剪枝

Hao Wu, Xudong Wang, Jialiang Zhang, Junlong Tong, Xinghao Chen, Junyan Lin, Yunpu Ma, Xiaoyu Shen

机构 * Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Shanghai Jiao Tong University(上海交通大学) The Hong Kong Polytechnic University(香港理工大学) Munich Center for Machine Learning, LMU Munich(慕尼黑机器学习中心,慕尼黑大学)

AI总结 UTPTrack通过统一令牌剪枝框架,在视觉跟踪中实现高精度与高效能的平衡,同时支持多模态和语言引导任务。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23645 2026-03-02 cs.CV

BuildAnyPoint: 3D Building Structured Abstraction from Diverse Point Clouds

BuildAnyPoint: 从多样点云构建三维建筑结构抽象

Tongyan Hua, Haoran Gong, Yuan Liu, Di Wang, Ying-Cong Chen, Wufan Zhao

机构 * HKUST(GZ)(香港科技大学(广州)) Xi’an Jiaotong University(西安交通大学)

AI总结 BuildAnyPoint通过Loca-DiT框架从多样点云中生成结构化三维建筑抽象,提升重建的表面准确性和分布均匀性。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23618 2026-03-02 cs.CV

Egocentric Visibility-Aware Human Pose Estimation

视角感知的人体姿态估计

Peng Dai, Yu Zhang, Yiqiang Feng, Zhen Fan, Yang Zhang

机构 * PICO

AI总结 本文提出Eva-3M数据集和EvaPose方法,通过引入关键点可见性标注提升视角感知人体姿态估计的准确性。

Comments Conference on Computer Vision and Pattern Recognition 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12769 2026-03-02 cs.CV

PixelRush: Ultra-Fast, Training-Free High-Resolution Image Generation via One-step Diffusion

PixelRush: 通过一步扩散实现超快、无训练高分辨率图像生成

Hong-Phuc Lai, Phong Nguyen, Anh Tran

机构 * Qualcomm AI Research(高通人工智能研究)

AI总结 PixelRush通过一步扩散实现超快无训练高分辨率图像生成,生成4K图像速度提升10-35倍,保持高质量输出。

Comments Accepted to CVPR 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00445 2026-03-02 cs.CV

AutoDebias: Automated Framework for Debiasing Text-to-Image Models

AutoDebias:文本到图像模型的自动化去偏框架

Hongyi Cai, Mohammad Mahdinur Rahman, Mingkang Dong, Muxin Pu, Moqyad Alqaily, Jie Li, Xinfeng Li, Jialie Shen, Meikang Qiu, Qingsong Wen

机构 * Universiti Malaya(马来大学) Monash University(莫纳什大学) United Arab Emirates University(阿拉伯联合酋长国大学) University of Science and Technology Beijing(北京科技大学) Nanyang Technological University (NTU)(南洋理工大学) City St George’s, University of London(伦敦大学圣乔治学院) Augusta University(奥古斯塔大学) Squirrel Ai Learning

AI总结 AutoDebias提出了一种自动化框架,用于检测和减轻文本到图像模型中的恶意偏见,通过视觉语言模型和CLIP引导的训练过程,有效应对隐秘注入的刻板印象和多重攻击。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18496 2026-03-02 cs.CV

Distilling Balanced Knowledge from a Biased Teacher

从偏见教师中提炼平衡知识

Seonghak Kim

机构 * Agency for Defense Development (ADD)(国防开发机构)

AI总结 LTKD通过引入跨组和组内损失,有效缓解长尾分布中教师模型的偏见,提升模型在尾部类别上的表现。

Comments 10 pages, 5 figures, accepted by The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16930 2026-03-02 cs.CV

What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?

什么使零样本立体匹配具有良好的合成训练数据?

David Yan, Alexander Raistrick, Jia Deng

机构 * Princeton University(普林斯顿大学)

AI总结 本文研究了合成数据集对零样本立体匹配性能的影响,通过参数调整和大规模数据集验证,发现其性能优于混合数据集并具有开源代码支持。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23575 2026-03-02 cs.CV cs.AI

CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird's-Eye-View Semantic Segmentation

CycleBEV: 通过视图循环一致性正则化视图转换网络用于鸟瞰图语义分割

Jeongbin Hong, Dooseop Choi, Taeg-Hyun An, Kyounghwan An, Kyoung-Wook Min

机构 * Electronics and Telecommunications Research Institute (ETRI)(电子电信研究院) University of Science and Technology (UST)(科学技术大学)

AI总结 CycleBEV通过引入视图循环一致性正则化方法,提升视图转换网络在鸟瞰图语义分割中的性能,实现显著的mIoU提升。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23559 2026-03-02 cs.CV

No Calibration, No Depth, No Problem: Cross-Sensor View Synthesis with 3D Consistency

无需校准,无需深度,无需问题:跨传感器视角合成与3D一致性

Cho-Ying Wu, Zixun Huang, Xinyu Huang, Liu Ren

机构 * Bosch Research North America(博世北美研究部) Bosch Center for Artificial Intelligence (BCAI)(博世人工智能中心)

AI总结 本文提出了一种无需校准的跨传感器视角合成方法,通过匹配、密集化和3D点云整合技术,解决RGB-X数据对齐问题,提升大规模真实世界数据收集的效率。

Comments CVPR 2026 Main Conference. Project page: https://choyingw.github.io/3d-rgbx.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23523 2026-03-02 cs.CV

All in One: Unifying Deepfake Detection, Tampering Localization, and Source Tracing with a Robust Landmark-Identity Watermark

三位一体:统一深度伪造检测、篡改定位和来源追溯的稳健地标-身份水印

Junjiang Wu, Liejun Wang, Zhiqing Guo

机构 * School of Computer Science and Technology, Xinjiang University, Urumqi, China(新疆大学计算机科学与技术学院) Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center, Urumqi, China(新疆多模态智能处理与信息安全工程技术研究中心)

AI总结 本文提出LIDMark框架,通过统一的地标-身份水印实现深度伪造检测、篡改定位和来源追溯的三位一体解决方案。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21399 2026-03-02 cs.LG cs.AI cs.CV

FedVG: Gradient-Guided Aggregation for Enhanced Federated Learning

FedVG: 基于梯度的聚合以提升联邦学习

Alina Devkota, Jacob Thrasher, Donald Adjeroh, Binod Bhattarai, Prashnna K. Gyawali

机构 * West Virginia University(西弗吉尼亚大学) University of Aberdeen(阿伯丁大学) Fogsphere (Redev.AI)(Fogsphere(Redev.AI))

AI总结 FedVG通过利用全局验证集指导梯度优化,提升联邦学习在异质数据环境下的泛化性能和聚合效果。

Comments Accepted to CVPR 2026 (Findings Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00805 2026-03-02 cs.CV

Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding

通过草稿思考:用于高效长视频理解的推测性时间推理

Pengfei Hu, Meng Cao, Yingyao Wang, Yi Wang, Jiahua Dong, Jun Song, Yu Cheng, Bo Zheng, Xiaodan Liang

机构 * MBZUAI(马克斯·普朗克智能研究院) Alibaba Group Holding Limited(阿里巴巴集团控股有限公司) Shanghai AI Lab(上海人工智能实验室)

AI总结 SpecTemp通过协作双模型设计,实现高效长视频理解,平衡效率与准确性,提升推理速度。

Comments Accepted by CVPR 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23361 2026-02-27 cs.CV

VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scale

VGG-T³:大规模离线前馈3D重建

Sven Elflein, Ruilong Li, Sérgio Agostinho, Zan Gojcic, Laura Leal-Taixé, Qunjie Zhou, Aljosa Osep

机构 * NVIDIA Vector Institute(向量研究所) University of Toronto(多伦多大学)

AI总结 VGG-T³通过测试时训练将变长键值空间蒸馏为固定大小的MLP,实现大规模离线前馈3D重建,速度比基线方法快11.6倍,并在重建精度和视觉定位能力上表现优异。

Comments CVPR 2026, Project page: https://research.nvidia.com/labs/dvl/projects/vgg-ttt

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23359 2026-02-27 cs.CV cs.AI

SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation

SeeThrough3D: 3D布局生成中的遮挡感知控制

Vaibhav Agrawal, Rishubh Parihar, Pradhaan Bhat, Ravi Kiran Sarvadevabhatla, R. Venkatesh Babu

机构 * IIIT Hyderabad(IIIT海得拉尔) IISc Bengaluru(IISc班加罗尔)

AI总结 SeeThrough3D通过引入遮挡感知的3D场景表示和掩码自注意力机制,实现了对3D布局生成中遮挡关系的精确建模与控制。

Comments Project page: https://seethrough3d.github.io. Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏