arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

至 收录 2109
2602.00637 2026-02-03 cs.CV

VIZOR: Viewpoint-Invariant Zero-Shot Scene Graph Generation for 3D Scene Reasoning

VIZOR: 视点不变的零样本场景图生成用于3D场景推理

Vivek Madhavaram, Vartika Sengar, Arkadipta De, Charu Sharma

机构 * Machine Learning Lab, IIIT Hyderabad, India(IIIT海得拉巴机器学习实验室) Fujitsu Research India, Bangalore(印度班加罗尔 Fujitsu 研究院)

AI总结 VIZOR通过无需训练的端到端框架,生成视点不变的零样本场景图,提升3D场景推理的准确性和泛化能力。

Comments WACV 2026, Project page: https://vivekmadhavaram.github.io/vizor/

URL PDF HTML 收藏
2511.18537 2026-02-03 cs.CV

Zero-Shot Video Deraining with Video Diffusion Models

无监督视频去雨与视频扩散模型

Tuomas Varanka, Juan Luis Gonzalez, Hyeongwoo Kim, Pablo Garrido, Xu Yao

机构 * University of Oulu(奥卢大学) Flawless AI Imperial College London(伦敦帝国理工学院)

AI总结 本文提出了一种无需合成数据和模型微调的无监督视频去雨方法,通过预训练文本到视频扩散模型,利用注意力切换机制提升动态场景中的去雨效果。

Comments WACV 2026

URL PDF HTML 收藏
2602.00109 2026-02-03 cs.CV eess.IV

Robustness of Presentation Attack Detection in Remote Identity Validation Scenarios

远程身份验证场景中呈现攻击检测的鲁棒性

John J. Howard, Richard O. Plesh, Yevgeniy B. Sirotin, Jerry L. Tipton, Arun R. Vemury

AI总结 本文研究了低光和自动图像采集对远程身份验证中PAD系统鲁棒性的影响,发现错误率显著增加,仅有一系统保持低错误率。

Comments Accepted to the IEEE/CVF WACV 2026 Workshop on Generative, Adversarial and Presentation Attacks in Biometrics (GAPBio). 8 pages, 6 figures, 4 tables

URL PDF HTML 收藏
2412.00112 2026-01-30 cs.CV cs.GR

BiPO: Bidirectional Partial Occlusion Network for Text-to-Motion Synthesis

BiPO:用于文本到动作合成的双向部分遮挡网络

Seong-Eun Hong, Soobin Lim, Juyeong Hwang, Minwook Chang, Hyeongyeop Kang

机构 * Kyung Hee University(庆尚大学) NC Research, NCSOFT Corp.(NC研究,NCSOFT公司) Korea University(韩国大学)

AI总结 BiPO通过整合部分生成与双向自回归架构,提升了文本到动作合成的性能,并在HumanML3D数据集上实现了最先进的结果。

Comments 18 pages, 11 figures. Accepted to WACV 2026 (Oral)

URL PDF HTML 收藏
2601.20661 2026-01-29 cs.CV

ProSkill: Segment-Level Skill Assessment in Procedural Videos

ProSkill:在程序化视频中进行段级技能评估

Michele Mazzamuto, Daniele Di Mauro, Gianpiero Francesca, Giovanni Maria Farinella, Antonino Furnari

机构 * University of Catania(卡塔尼亚大学) Next Vision s.r.l.(Next Vision公司) Toyota Motor Europe(丰田欧洲公司)

AI总结 ProSkill是首个针对程序化任务动作级技能评估的基准数据集,通过新颖的注释协议提供绝对和成对的技能评估,用于评估最新技能评估算法。

Comments Accepted at The IEEE/CVF Winter Conference on Applications of Computer Vision 2026

URL PDF HTML 收藏
2411.00639 2026-01-29 cs.CV

Event-guided Low-light Video Semantic Segmentation

事件引导的低光照视频语义分割

Zhen Yao, Mooi Choo Chuah

机构 * Lehigh University(莱文斯顿大学)

AI总结 本文提出EVSNet,利用事件模态引导学习光照不变表示,实现低光照环境下视频语义分割的高效高精度

Comments 12 pages, 5 figures, Accepted to IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2025

URL PDF HTML 收藏
2510.03906 2026-01-27 cs.CV

From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance

从滤波器到视觉语言模型:通过目标检测和分割性能评估去雾方法

Ardalan Aryashad, Parsa Razmara, Amin Mahjoub, Seyedarmin Azizi, Mahdi Salmani, Arad Firouzkouhi

机构 * University of Southern California(南加州大学)

AI总结 本文通过目标检测和分割性能评估,探讨了去雾方法在真实与合成环境中的有效性,揭示了视觉语言模型在恶劣天气下的应用潜力。

Comments Accepted at WACV 2026 Proceedings (Oral), 5th Workshop on Image, Video, and Audio Quality Assessment in Computer Vision, with a focus on VLM and Diffusion Models

URL PDF HTML 收藏
2410.05270 2026-01-27 cs.CV

CLIP's Visual Embedding Projector is a Few-shot Cornucopia

CLIP的视觉嵌入投影器是少样本宝库

Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick Pérez, Raoul de Charette

机构 * Inria(法国国家信息与自动化技术研究所) Kyutai

AI总结 ProLIP通过正则化投影矩阵提升CLIP在少样本分类及跨领域迁移中的性能,提供更高效的替代方案。

Comments WACV 2026

URL PDF HTML 收藏
2601.18001 2026-01-27 cs.CV

MorphXAI: An Explainable Framework for Morphological Analysis of Parasites in Blood Smear Images

MorphXAI: 一种用于血涂片图像中寄生虫形态分析的可解释框架

Aqsa Yousaf, Sint Sint Win, Megan Coffee, Habeeb Olufowobi

机构 * Department of Computer Science and Engineering, University of Texas at Arlington(德克萨斯大学阿灵顿分校计算机科学与工程系) Department of Medicine, Division of Infectious Diseases, NYU Grossman School of Medicine(纽约大学格罗斯曼医学院医学系感染病科)

AI总结 MorphXAI通过整合形态学监督,实现了血涂片中寄生虫的检测与细粒度形态分析,提供结构化生物意义的解释。

Comments Accepted at WACV 2026

Journal ref Winter Conference on Applications of Computer Vision 2026

URL PDF HTML 收藏
2601.17927 2026-01-27 cs.CV cs.MM

RemEdit: Efficient Diffusion Editing with Riemannian Geometry

RemEdit: 基于黎曼几何的高效扩散编辑

Eashan Adhikarla, Brian D. Davison

AI总结 RemEdit通过基于黎曼几何的潜在空间导航和任务特定注意力剪枝机制,实现了高效且保真的图像编辑,同时保持实时性能。

Journal ref IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026

URL PDF HTML 收藏
2502.00662 2026-01-27 cs.CV cs.CL cs.LG

Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias Estimation

弥合模态差距:基于多模态原型和图像偏差估计的少样本分布外检测

Yimu Wang, Evelien Riddell, Adrian Chow, Sean Sedwards, Krzysztof Czarnecki

机构 * University of Waterloo(滑铁卢大学)

AI总结 本文提出SUPREME框架,通过引入多模态原型和图像偏差估计,有效缓解图像与文本之间的模态差距,提升少样本分布外检测性能。

Comments WACV 2026

URL PDF HTML 收藏
2408.09650 2026-01-27 cs.CV cs.AI cs.MM eess.IV

From Darkness to Detail: Frequency-Aware SSMs for Low-Light Vision

从黑暗到细节:面向低光视觉的频率感知SSM

Eashan Adhikarla, Kai Zhang, Gong Chen, John Nicholson, Brian D. Davison

机构 * Lehigh University(莱特大学) Lenovo Research(联想研究院)

AI总结 ExpoMamba通过频率感知状态空间模型解决低光图像增强中的混合曝光问题,实现高效高质量的实时增强。

Journal ref Winter Conference on Applications of Computer Vision, WACV 2026

URL PDF HTML 收藏
2601.16645 2026-01-26 cs.CV

Edge-Aware Image Manipulation via Diffusion Models with a Novel Structure-Preservation Loss

通过扩散模型实现边缘感知图像处理:一种新的结构保持损失

Minsu Gong, Nuri Ryu, Jungseul Ok, Sunghyun Cho

机构 * Planby Technologies(Planby技术公司) POSTECH

AI总结 本文提出了一种新的结构保持损失,用于提升潜在扩散模型在图像编辑中的边缘结构保真度,通过局部线性模型量化结构差异并结合后处理步骤和掩码策略,实现了更高质量的图像编辑效果。

Comments Accepted to WACV 2026

URL PDF HTML 收藏
2507.00243 2026-01-26 cs.CV

VOCAL: Visual Odometry via ContrAstive Learning

基于对比学习的视觉里程计:VOCAL

Chi-Yao Huang, Zeel Bhatt, Yezhou Yang

机构 * Arizona State University(亚利桑那州立大学)

AI总结 VOCAL通过对比学习将视觉里程计重新定义为标签排序问题,提升可解释性和灵活性,推动更通用的空间智能发展。

Comments Accepted to WACV 2026

URL PDF HTML 收藏
2409.11355 2026-01-23 cs.CV

Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think

对图像条件扩散模型进行微调比你想象的更容易

Gonzalo Martin Garcia, Karim Knaebel, Christian Schmidt, Daan de Geus, Alexander Hermans, Bastian Leibe

机构 * RWTH Aachen University(亚琛工业大学) Eindhoven University of Technology(埃因霍温理工大学)

AI总结 本文发现对图像条件扩散模型进行微调比预期更简单,并展示了其在深度和法线估计任务中的优越性能。

Comments WACV 2025 Oral. Project page at https://vision.rwth-aachen.de/diffusion-e2e-ft

URL PDF HTML 收藏
2601.15711 2026-01-23 cs.CV

Zero-Shot Product Attribute Labeling with Vision-Language Models: A Three-Tier Evaluation Framework

基于视觉-语言模型的零样本产品属性标注:一种三级评估框架

Shubham Shukla, Kunal Sonalkar

机构 * Nordstrom(诺德斯特拉姆)

AI总结 本文提出一种三级评估框架,评估视觉-语言模型在多属性服装任务中的零样本属性标注性能,发现高效模型在成本较低时能实现接近旗舰级性能,揭示了属性适用性检测是关键瓶颈。

Comments Accepted to WACV 2026 Workshop on Physical Retail AI (PRAW)

URL PDF HTML 收藏
2509.25856 2026-01-23 cs.CV

PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection

PatchEAD: 统一工业视觉提示框架用于补丁专属异常检测

Po-Han Huang, Jeng-Lin Li, Po-Hsuan Huang, Ming-Ching Chang, Wei-Chao Chen

机构 * Inventec Corporation(Inventec公司) University at Albany, State University of New York(纽约州立大学阿尔巴尼分校)

AI总结 PatchEAD提出统一的补丁聚焦框架,实现无需训练的工业异常检测,兼容多种基础模型并提升补丁相似性鲁棒性。

Comments 10 pages, 5 figures. WACV 2026 (Accepted)

URL PDF HTML 收藏
2601.14154 2026-01-21 cs.CV cs.AI

LLM Augmented Intervenable Multimodal Adaptor for Post-operative Complication Prediction in Lung Cancer Surgery

基于大语言模型的可干预多模态适配器用于肺癌手术后并发症预测

Shubham Pandey, Bhavin Jawade, Srirangaraj Setlur, Venu Govindaraju, Kenneth Seastedt

机构 * University at Buffalo(布法罗大学) Roswell Park Comprehensive Cancer Center(罗斯威尔帕克综合癌症中心)

AI总结 MIRACLE通过整合术前临床和放射学数据,利用超球面嵌入空间和干预式深度学习模块,实现肺癌手术后并发症风险的预测与可解释性管理。

Comments Accepted to P2P-CV @ WACV 2026

URL PDF HTML 收藏
2601.14038 2026-01-21 cs.CV

Correcting and Quantifying Systematic Errors in 3D Box Annotations for Autonomous Driving

校正并量化自动驾驶中3D盒标注的系统误差

Alexandre Justo Miro, Ludvig af Klinteberg, Bogdan Timus, Aron Asefaw, Ajinkya Khoche, Thomas Gustafsson, Sina Sharif Mansouri, Masoud Daneshtalab

机构 * Traton Group R&D(特龙集团研发部) Mälardalen University(马尔默达伦大学) KTH Royal Institute of Technology(皇家理工学院)

AI总结 本研究提出了一种方法,用于校正和量化自动驾驶中3D盒标注的系统误差,提升标注质量并提高性能评估的准确性。

Comments Accepted to The IEEE/CVF Winter Conference on Applications of Computer Vision 2026

URL PDF HTML 收藏
2601.13974 2026-01-21 cs.CV

STEC: A Reference-Free Spatio-Temporal Entropy Coverage Metric for Evaluating Sampled Video Frames

STEC:一种无参考的时空熵覆盖度量,用于评估采样视频帧

Shih-Yao Lin

机构 * Independent Researcher(独立研究者)

AI总结 STEC是一种无参考的时空熵覆盖度量,用于评估视频帧采样的有效性,通过联合建模空间信息强度、时间分散性和非冗余性,提供一种轻量且原则性的采样质量度量。

Comments This paper corresponds to the camera-ready version of a WACV 2026 Workshop paper

URL PDF HTML 收藏
2601.13502 2026-01-21 cs.CV

DIS2: Disentanglement Meets Distillation with Classwise Attention for Robust Remote Sensing Segmentation under Missing Modalities

DIS2: 通过类级注意机制实现解耦与蒸馏以在缺失模态下实现鲁棒的遥感分割

Nhi Kieu, Kien Nguyen, Arnold Wiliem, Clinton Fookes, Sridha Sridharan

机构 * Queensland University of Technology(昆士兰理工大学) Shield AI

AI总结 DIS2通过类级注意机制实现解耦与蒸馏,以在缺失模态下实现鲁棒的遥感分割。

Comments Accepted to WACV 2026 - Computer Vision for Earth Observation Workshop

URL PDF HTML 收藏
2601.12814 2026-01-21 cs.CV

CSGaussian: Progressive Rate-Distortion Compression and Segmentation for 3D Gaussian Splatting

CSGaussian:渐进率失真压缩与分割用于3D高斯溅射

Yu-Jen Tseng, Chia-Hao Kao, Jing-Zhong Chen, Alessandro Gnutti, Shao-Yuan Lo, Yen-Yu Lin, Wen-Hsiao Peng

机构 * National Yang Ming Chiao Tung University University of Brescia National Taiwan University

AI总结 CSGaussian提出一种统一框架,通过整合语义学习实现3D高斯溅射的率失真优化压缩与分割,提升场景编辑与操作能力。

Comments Accepted at WACV 2026

URL PDF HTML 收藏
2601.12636 2026-01-21 cs.CV

From Bands to Depth: Understanding Bathymetry Decisions on Sentinel-2

从带宽到深度:理解Sentinel-2的水深决策

Satyaki Roy Chowdhury, Aswathnarayan Radhakrishnan, Hsiao Jou Hsu, Hari Subramoni, Joachim Moortgat

机构 * The Ohio State University(俄亥俄州立大学)

AI总结 本文提出Swin-BathyUNet模型,通过分析其深度推断机制和可靠性,揭示了Sentinel-2水深决策中光谱重要性、注意力机制及跨区域推理的挑战与解决方案。

Comments Accepted by WACV 2026

URL PDF HTML 收藏
2601.12493 2026-01-21 cs.CV

Histopath-C: Towards Realistic Domain Shifts for Histopathology Vision-Language Adaptation

Histopath-C: 向 histopathology 视觉-语言适应的现实领域偏移迈进

Mehrdad Noori, Gustavo Adolfo Vargas Hakim, David Osowiechi, Fereshteh Shakeri, Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Ismail Ben Ayed, Christian Desrosiers

机构 * LIVIA, ÉTS Montreal, Canada International Laboratory on Learning Systems (ILLS)(LIVIA,蒙特利尔工程学院,加拿大国际学习系统实验室)

AI总结 Histopath-C通过引入现实合成损坏的基准和LATTE策略,提升组织病理学图像的视觉-语言适应鲁棒性。

Comments Accepted to WACV 2026

URL PDF HTML 收藏
2508.02927 2026-01-21 cs.CV

Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?

红外小规模卷积网络目标检测:ImageNet预训练仍然有用吗?

Srikanth Muralidharan, Heitor R. Medeiros, Masih Aminbeidokhti, Eric Granger, Marco Pedersoli

机构 * LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(LIVIA,系统工程系,蒙特利尔大学) International Laboratory on Learning Systems (ILLS)(学习系统国际实验室)

AI总结 本文研究了超小型卷积网络在红外目标检测中的性能,发现ImageNet预训练在一定容量阈值后对鲁棒性提升有限,建议避免过于小的模型以提高鲁棒性。

Comments Accepted to WACV 2026

URL PDF HTML 收藏
2507.08711 2026-01-21 cs.CV

SGPMIL: Sparse Gaussian Process Multiple Instance Learning

SGPMIL:稀疏高斯过程多实例学习

Andreas Lolos, Stergios Christodoulidis, Aris L. Moustakas, Jose Dolz, Maria Vakalopoulou

机构 * National and Kapodistrian University of Athens(希腊国家与卡波迪斯蒂亚诺斯大学) ÉTS Montréal(蒙特利尔ÉTS) Archimedes, Athena Research Center(阿提卡研究中心) CentraleSupélec, Université Paris-Saclay(巴黎萨克雷大学中央理工-supélec)

AI总结 SGPMIL通过引入稀疏高斯过程,提升多实例学习中实例级预测的可靠性与可解释性。

Comments 8 pages, 4 figures, 2 tables. Accepted to IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2601.10921 2026-01-19 cs.CV cs.AI cs.LG

RobuMTL: Enhancing Multi-Task Learning Robustness Against Weather Conditions

RobuMTL: 提高多任务学习对天气条件的鲁棒性

Tasneem Shaffee, Sherief Reda

机构 * Brown University(布朗大学)

AI总结 RobuMTL通过动态选择LoRA模块提升多任务学习在恶劣天气下的鲁棒性,实验证明在PASCAL和NYUD-v2数据集上均取得显著性能提升。

Comments Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2601.10836 2026-01-19 cs.CV

One Model, Many Behaviors: Training-Induced Effects on Out-of-Distribution Detection

一个模型,多种行为:训练诱导对分布外检测的影响

Gerhard Krumpl, Henning Avenhaus, Horst Possegger

机构 * Institute of Visual Computing, Graz University of Technology(视觉计算研究所,格拉茨技术大学) KESTRELEYE GmbH(KESTRELEYE公司)

AI总结 本文研究了训练对分布外检测性能的影响,发现准确性与检测性能之间存在非单调关系,并揭示了训练策略与检测方法间的强相关性。

Comments WACV 2026

URL PDF HTML 收藏
2601.10802 2026-01-19 cs.CV

ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research

ICONIC-444: 一个310万张图像数据集用于OD检测研究

Gerhard Krumpl, Henning Avenhaus, Horst Possegger

机构 * Institute of Visual Computing, Graz University of Technology(视觉计算研究所,格拉茨技术大学) KESTRELEYE GmbH(KESTRELEYE公司)

AI总结 ICONIC-444是一个包含310万张图像的工业数据集,旨在通过提供多样化的数据支持OD检测研究,涵盖444个类别并定义四个参考任务以评估不同方法的性能。

Comments WACV 2026, Dataset repo: https://github.com/gkrumpl/iconic-444

URL PDF HTML 收藏
2506.12198 2026-01-19 cs.CV

ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models

ViSTA: 基于多模态适配器的文本到图像扩散模型用于视觉叙事

Sibo Dong, Ismail Shaheen, Maggie Shen, Rupayan Mallick, Sarah Adel Bargal

机构 * Department of Computer Science, Georgetown University(计算机科学系,乔治城大学)

AI总结 ViSTA通过多模态历史适配器提升文本到图像扩散模型在视觉叙事中的生成一致性与文本对齐能力。

Comments Accepted to WACV 2026

URL PDF HTML 收藏