arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

至 收录 2109
2608.13007 2026-08-14 cs.CV 新提交

Structure-aware Riemannian Growth Fields for 4D Plant Modeling

面向4D植物建模的结构感知黎曼生长场

Meng-Yu Jennifer Kuo, Ryo Kawahara

机构 * Nara Women’s University(奈良女子大学) Kyoto University(京都大学)

AI总结 该研究提出结构感知黎曼生长场框架,用于从稀疏时间观测重建4D植物生长,构建了10天双物种标注数据集,在几何精度与对应一致性上优于现有方法。

Comments Accepted to WACV 2027 (Round 1)

URL PDF HTML 收藏
2511.17361 2026-08-12 cs.CV 版本更新

SuperQuadricOcc: Real-Time Self-Supervised Semantic Occupancy Estimation with Superquadric Volume Rendering

SuperQuadricOcc: 基于超级二次曲面体渲染的实时自监督语义占用估计

Seamie Hayes, Alexandre Boulch, Andrei Bursuc, Reenu Mohandas, Ganesh Sistu, Tim Brophy, Ciaran Eising

机构 * Data Driven Computer Engineering (D²iCE) Research Centre(数据驱动计算机工程(D²iCE)研究中心) University of Limerick(利默里克大学) Taighde Éireann – Research Ireland(爱尔兰研究——塔吉德)

AI总结 本文提出SuperQuadricOcc,通过超级二次曲面体渲染实现实时自监督语义占用估计,相比传统方法更高效且占用内存更少,在Occ3D-nuScenes数据集上取得最佳性能。

Comments Accepted at WACV 2027

URL PDF HTML 收藏
2608.07632 2026-08-11 q-bio.QM cs.CV eess.IV 新提交

JUMP-lite: Compact, reproducible benchmarking of cell representations

JUMP-lite:紧凑、可复现的细胞表征基准测试

Alán F. Muñoz, Johan Fredin Haslum, Runxi Shen, Anne E. Carpenter, Shantanu Singh

AI总结 研究针对JUMP数据集体积过大导致细胞表征基准测试难以开展的问题,提出JUMP-lite压缩子集与Nahual框架,测试5种表征方法并验证压缩保留下游信号,为相关基准测试提供基础。

Comments Submitted to WACV 2027

URL PDF HTML 收藏
2608.09579 2026-08-11 cs.CV 新提交

You Only Flow Once: Calibrated and Real-Time Radar Pose Estimation with Multi-Hypothesis Normalizing Flows

仅一次流:基于多假设归一化流的校准且实时雷达位姿估计

Jonas Leo Mueller, Sebastian Hoefler, Dario Zanca, Naga Venkata Sai Jitin Jami, Thomas Altstidl, Bjoern M. Eskofier

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔兰根-纽伦堡大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) LMU München(慕尼黑大学) Helmholtz Zentrum München(慕尼黑亥姆霍兹中心)

AI总结 该研究针对雷达位姿估计的歧义问题,提出MH-NFPG方法,结合时空Transformer与归一化流,实现高效实时的校准位姿估计,性能优于扩散模型。

Comments Accepted at the Winter Conference on Applications of Computer Vision (WACV) 2027

URL PDF HTML 收藏
2608.07559 2026-08-11 cs.CV cs.LG 新提交

MVMD: A Multi-View Approach for Enhanced Mirror Detection

MVMD:用于增强镜面检测的多视图方法

Yidan Shen, Yu Wen, Chen Zhang, Xin Fu, Renjie Hu

机构 * University of Houston(休斯顿大学)

AI总结 针对现有镜面检测仅关注单图像的局限,提出多视图镜面检测方法MVMD及首个多视图镜面检测数据库,通过三个模块提升检测效果,使准确率和IoU分别最高提升2.6%和11.1%,增强了镜面密集环境下的三维重建准确性。

Comments This work has already published at WACV 2025, just want more accessibility

Journal ref 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

URL PDF HTML 收藏
2601.02927 2026-08-11 cs.CV cs.AI

PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding

PrismVAU: 用于多模态视频异常理解的提示优化推理系统

Iñaki Erregue, Kamal Nasrollahi, Sergio Escalera

机构 * Universitat de Barcelona(巴塞罗那大学) Computer Vision Center(计算机视觉中心) Aalborg University(奥胡斯大学) Milestone Systems(Milestone系统)

AI总结 PrismVAU通过轻量级系统和自动提示工程实现高效的多模态视频异常理解,无需复杂标注和外部模块。

Comments This paper has been accepted to the 6th Workshop on Real-World Surveillance: Applications and Challenges (WACV 2026)

URL PDF HTML 收藏
2601.04824 2026-08-07 cs.CV

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

SOVABench:多模态大语言模型的车辆监控动作检索基准

Oriol Rabasseda, Zenjie Li, Kamal Nasrollahi, Sergio Escalera

机构 * Milestone Systems A/S(Milestone Systems公司) Universitat de Barcelona(巴塞罗那大学) Computer Vision Center(计算机视觉中心) Aalborg Universitet(奥胡斯大学)

AI总结 SOVABench为多模态大语言模型提供车辆监控动作检索基准,通过定义两种评估协议评估跨动作区分和时间方向理解,展示了模型在复杂监控任务中的性能。

Comments This work has been accepted at Real World Surveillance: Applications and Challenges, 6th (in WACV Workshops)

URL PDF HTML 收藏
2501.14198 2026-08-07 eess.IV cs.CV 版本更新

Sparse Mixture-of-Experts for Non-Uniform Noise Reduction in MRI Images

用于MRI图像非均匀降噪的稀疏混合专家模型

Zeyun Deng, Joseph Campbell

机构 * Purdue University(普渡大学)

AI总结 针对MRI图像非均匀噪声问题,提出细粒度稀疏混合专家框架,将图像分区域后路由至专用降噪CNN,在合成与真实脑部MRI数据集上性能优于现有方法,且泛化性良好。

Comments Accepted to the WACV Workshop on Image Quality

Journal ref in 2025 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW), Tucson AZ USA, pp 260-268

URL PDF HTML 收藏
2406.17109 2026-08-05 cs.CV

GMT: Guided Mask Transformer for Leaf Instance Segmentation

Feng Chen, Sotirios A. Tsaftaris, Mario Valerio Giuffrida

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2025 (Oral Presentation)

Journal ref IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2025, pp. 1217-1226

URL PDF HTML 收藏
2608.02309 2026-08-04 cs.CV 新提交

CalibBEV: LiDAR-Camera Calibration via BEV Alignment

CalibBEV:基于鸟瞰图对齐的激光雷达-相机标定方法

Filippo D'Addeo, Lorenzo Cipelli, Adriano Cardace, Emanuele Ghelfi, Andrea Zinelli, Massimo Bertozzi

机构 * University of Bologna(博洛尼亚大学) University of Parma(帕尔马大学) Stanford University(斯坦福大学) VisLab srl(VisLab有限公司) Ambarella Inc.(安霸公司)

AI总结 CalibBEV是一种激光雷达-相机标定方法,通过两步BEV对齐结合CLIP对比损失实现跨模态统一,在KITTI、nuScenes基准上大幅降低了标定误差,达到最优性能。

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2026. p. 4345-4354

URL PDF HTML 收藏
2504.05615 2026-07-30 cs.LG cs.AI

FedEFC: Federated Learning Using Enhanced Forward Correction Against Noisy Labels

FedEFC: 在噪声标签下使用增强前向校正的联邦学习

Seunghun Yu, Jin-Hyun Ahn, Joonhyuk Kang

机构 * KAIST(韩国科学技术院) Myongji University(明溪大学)

AI总结 FedEFC通过预停止和损失校正技术,有效缓解联邦学习中噪声标签的影响,尤其在异构数据环境下表现优异。

Comments 9 pages, 3 figures, revised version

Journal ref Proc. IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

URL PDF HTML 收藏
2607.22702 2026-07-28 cs.CV cs.LG 新提交

MIME: Multimodal Interactive Motion Encoder

MIME:多模态交互式运动编码器

Addison Zucek, Prerit Gupta, Kamila Kuatova, Aniket Bera

机构 * Purdue University(普渡大学)

AI总结 研究针对动画、AR/VR等场景中多人交互的文本-运动表示学习问题,提出多模态交互式运动编码器MIME,通过特定方法捕捉结构并训练,在文本-运动检索任务中表现出色,还能跨数据集支持下游运动生成。

Comments Under review at WACV 2027

URL PDF HTML 收藏
2409.06067 2026-07-08 cs.AI cs.CL cs.LG 版本更新

MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning

MLLM-LLaVA-FL:多模态大语言模型辅助联邦学习

Jianyi Zhang, Hao Frank Yang, Ang Li, Xin Guo, Pu Wang, Haiming Wang, Yiran Chen, Hai Li

机构 * Duke University(杜克大学) Johns Hopkins University(约翰霍普金斯大学) University of Maryland College Park(马里兰大学学院市分校) Lenovo Research(联想研究院)

AI总结 针对联邦学习中数据异质性问题,提出MLLM-LLaVA-FL框架,利用多模态大语言模型,通过全球视觉文本预训练、客户端本地训练和服务器端全局对齐三个阶段,提升联邦学习性能。

Comments Accepted to WACV 2025

Journal ref IEEE/CVF Winter Conference on Applications of Computer Vision (WACV 2025)

URL PDF HTML 收藏
2508.09629 2026-07-07 cs.CV 版本更新

Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors

利用学习到的纹理先验增强单目3D手部重建

Giorgos Karvounas, Nikolaos Kyriazis, Iason Oikonomidis, Georgios Pavlakos, Antonis A. Argyros

机构 * ICS-FORTH(希腊研究所) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Crete(克里特大学)

AI总结 重新审视纹理在单目3D手部重建中的作用,提出轻量级纹理模块,嵌入像素观察到UV纹理空间,实现预测与观察手部外观的密集对齐损失,增强HaMeR系统,提高精度和真实感。

Comments Accepted at WACV 2026. Project page: https://gkarv.github.io/hand-texture-module/

URL PDF HTML 收藏
2401.10805 2026-07-07 cs.CV cs.AI cs.LG cs.RO 版本更新

Learning to Visually Connect Actions and their Effects

学习在视觉上连接动作及其效果

Paritosh Parmar, Eric Peh, Basura Fernando

机构 * Institute of High-Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科技研究局,新加坡)

AI总结 该研究引入视觉连接动作及其效果(CATE)概念,探索其动作选择和效果亲和力评估两方面,设计基线模型,发现模型在该任务中表现不佳,人类远超模型,还表明CATE可作为自监督任务学习视频表征。

Comments WACV 2025 (Two Reviewer Nominations for Best Paper Candidate; Oral Presentation)

URL PDF HTML 收藏
2211.10872 2026-07-07 cs.CV 版本更新

MetaMax: Improved Open-Set Deep Neural Networks via Weibull Calibration

MetaMax:通过威布尔校准改进开放集深度神经网络

Zongyao Lyu, Nolan B. Gutierrez, William J. Beksi

机构 * The University of Texas at Arlington(德克萨斯大学阿灵顿分校)

AI总结 研究开放集识别问题,提出MetaMax这一有效后处理技术,直接对类激活向量建模,改进当代方法,无需计算类平均激活向量等,实验表明其性能优于OpenMax且与其他先进方法相当。

Comments To be presented at the 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshop on Dealing with Novelty in Open Worlds (DNOW); v2 added related work section

URL PDF HTML 收藏
2607.01621 2026-07-03 cs.AI 新提交

Spatial Support Matters: Geometry-Aware Graph Fusion for Rainfall Field Reconstruction

空间支撑至关重要:面向降雨场重建的几何感知图融合

Low Jun Yu, Niramay Kachhadiya, Herath Mudiyanselage Viraj Vidura Herath, Sanka Rasnayaka, Lucy Amanda Marshall

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) Faculty of Engineering, The University of Sydney(悉尼大学工程学院)

AI总结 针对降雨场重建中多源观测支撑(点、线、面)的几何差异问题,提出几何感知多支撑异构图神经网络,通过跨支撑消息传递融合并重建降雨场,在新加坡数据上RMSE降低23.2%。

Comments Submitted to WACV 2027, applications track

URL PDF HTML 收藏
2404.10034 2026-07-01 cs.CV cs.LG 版本更新

A Realistic Protocol for Evaluation of Weakly Supervised Object Localization

弱监督目标定位评估的现实协议

Shakeeb Murtaza, Soufiane Belharbi, Marco Pedersoli, Eric Granger

机构 * LIVIA, ILLS, Dept. of Systems Engineering, ETS Montreal(蒙特利尔高等技术学院系统工程系LIVIA实验室、ILLS实验室)

AI总结 提出一种无需手动边界框标注的WSOL评估协议,利用预训练区域提议方法生成伪边界框进行模型选择和阈值估计,在自然和医学图像数据集上达到与使用真实边界框相当的性能。

Comments 13 pages, 5 figures

Journal ref WACV 2025: IEEE/CVF Winter Conf. on Applications of Computer Vision, Arizona, USA

URL PDF HTML 收藏
2606.29964 2026-06-30 cs.CV

Variance Reduction on the Camera Axis: Multi-View Score Distillation for 3D

相机轴上的方差缩减:用于3D的多视角分数蒸馏

Marian Lupascu, Mihai Sorin Stupariu, Ionut Mironica

机构 * Department of Computer Science, University of Bucharest(布加勒斯特大学计算机科学系) Adobe Research(Adobe研究院)

AI总结 提出多视角聚合分数蒸馏(MV-SDI),通过在每个优化步骤聚合K个视角的梯度来降低方差,无需重新训练或使用多视角数据,在固定UNet预算下显著提升3D生成的一致性和对齐指标。

Comments 30 pages, 19 figures. Submitted to WACV 2027 (Algorithms Track)

URL PDF HTML 收藏
2512.02456 2026-06-30 cs.CV cs.CL

See, Think, Learn: A Self-Taught Multimodal Reasoner

看见、思考、学习:一种自教的多模态推理器

Sourabh Sharma, Sonam Gupta, Sadbhawna

AI总结 本文提出See-Think-Learn框架,通过自训练提升多模态推理能力,结合感知与推理生成结构化推理过程,增强模型区分正确与误导性回答的能力。

Comments Accepted at The Winter Conference on Applications of Computer Vision 2026

URL PDF HTML 收藏
2606.20241 2026-06-19 cs.CV 新提交

BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models

BAFIS:评估现代文本到图像模型中的职业偏见与人类偏好的数据集与框架

Thomas Klassert, Adrian Ulges, Biying Fu

机构 * RheinMain University of Applied Sciences(莱茵美因应用科学大学)

AI总结 本研究提出BAFIS平台和包含21,140张多语言提示生成图像的数据集,评估五种文本到图像模型在职业生成中的性别和种族偏见,结合人类偏好反馈,发现系统性偏见并强调纳入人类偏好的必要性。

Comments Accepted at the IEEE Winter Conference on Applications of Computer Vision, WACV 2026

URL PDF HTML 收藏
2411.16934 2026-06-18 cs.CV

Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory

在线事件记忆视觉查询定位与眼动流对象记忆

Zaira Manigrasso, Matteo Dunnhofer, Antonino Furnari, Moritz Nottebaum, Antonio Finocchiaro, Davide Marana, Rosario Forte, Giovanni Maria Farinella, Christian Micheloni

机构 * University of Udine(乌迪内大学) University of Catania(卡塔尼亚大学) York University(约克大学)

AI总结 本文提出OVQ2D任务,通过在线处理视频流实现对象定位,引入ESOM框架整合发现、跟踪和记忆模块,实验显示其在Ego4D数据集上表现优异,但仍有提升空间。

Comments in IEEE/CVF Winter Conference on Application of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2605.25921 2026-05-26 cs.GR cs.CV

Curve Skeletonization in Continuous domain for Meshes and Point Clouds

网格与点云的连续域曲线骨架化

Jai Bardhan, Ramya Hebbalaguppe, Aravind Udupa

机构 * TCS Research(TCS研究) IIT Delhi(德里理工学院)

AI总结 提出CSCD框架,将基于局部分隔符的骨架化方法推广到连续域,通过CSCD-M(网格)和CSCD-PC(点云)两种实现,提升了骨架提取的鲁棒性和拓扑保持能力。

Comments 31 pages, 26 figures, 7 tables, 4 algorithms. Published at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2601.08205 2026-05-26 cs.CV cs.LG

FUME: Fused Unified Multi-Gas Emission Network for Livestock Rumen Acidosis Detection

FUME: 用于牲畜瘤胃酸中毒检测的融合统一多气体排放网络

Taminul Islam, Toqi Tahamid Sarker, Mohamed Embaby, Khaled R Ahmed, Amer AbuGhazaleh

机构 * Southern Illinois University, Carbondale(南方伊利诺伊大学,卡本达勒分校) University of California, Davis(加州大学戴维斯分校)

AI总结 提出FUME网络,利用双气体(CO2和CH4)光学成像,通过轻量双流架构和通道注意力融合,实现瘤胃酸中毒的高精度分割与分类。

Comments 10 pages, 5 figures

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops, 2026, pp. 510-519

URL PDF HTML 收藏
2508.01014 2026-05-18 cs.RO cs.CV

Hestia: Voxel-Face-Aware Hierarchical Next-Best-View Acquisition for Efficient 3D Reconstruction

Hestia:面向高效3D重建的体素-面感知分层最佳视角获取

Cheng-You Lu, Zhuoli Zhuang, Nguyen Thanh Trung Le, Da Xiao, Yu-Cheng Chang, Thomas Do, Srinath Sridhar, Chin-teng Lin

机构 * University of Technology Sydney(悉尼技术大学) Brown University(布朗大学)

AI总结 本文提出Hestia,一种面向高效3D重建的体素-面感知分层最佳视角获取方法,通过改进的规划器组件提升鲁棒性和性能,实验显示在覆盖比和Chamfer距离上均有显著提升。

Comments Accepted to the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2407.12173 2026-05-12 cs.CV cs.AI

Beta Sampling is All You Need: Efficient Image Generation Strategy for Diffusion Models using Stepwise Spectral Analysis

贝塔采样是全部所需:利用分步频谱分析的扩散模型高效图像生成策略

Haeil Lee, Hansang Lee, Seoyeon Gye, Junmo Kim

机构 * School of Electrical Engineering, KAIST(韩国科学技术院电子工程学院)

AI总结 本文提出基于频谱分析的贝塔采样方法,优化扩散模型去噪过程,通过重点投入关键步骤提升生成效率与质量,实验显示其在FID和IS评分上优于传统均匀采样。

Comments 8 pages, 9 figures, WACV 2025

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 4215-4224, 2025

URL PDF HTML 收藏
2512.07703 2026-05-01 cs.CV cs.LG

PVeRA: Probabilistic Vector-Based Random Matrix Adaptation

PVeRA:基于概率向量的随机矩阵适应

Leo Fillioux, Enzo Ferrante, Paul-Henry Cournède, Maria Vakalopoulou, Stergios Christodoulidis

机构 * IHU PRISM, National Center for Precision Medicine in Oncology, Gustave Roussy(IHU PRISM,国家精准医学肿瘤中心,古斯塔夫·罗斯-伊实验室) Institute of Computer Sciences, CONICET, Universidad de Buenos Aires(计算机科学研究所,CONICET,布宜诺斯艾利斯大学)

AI总结 本文提出PVeRA,一种基于概率的随机矩阵适应方法,通过概率方式修改VeRA中的低秩矩阵,提升模型在输入模糊性和不同采样配置下的适应能力,实验显示其在VTAB-1k基准上优于其他适配器。

Journal ref WACV 2026

URL PDF HTML 收藏
2503.07878 2026-05-01 cs.CV cs.AI

A Woman with a Knife or A Knife with a Woman? Measuring Directional Bias Amplification in Image Captions

一个女人持刀还是一把刀持一个女人?测量图像描述中的方向性偏见放大

Rahul Nair, Bhanu Tokas, Hannah Kerner

机构 * Arizona State University(亚利桑那州立大学)

AI总结 本文提出DBAC指标,用于测量图像描述中偏见放大现象,改进了现有LIC方法,提升了对偏见来源的识别能力,并在COCO数据集上验证了其有效性。

Comments Accepted at WACV 2026. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2026

URL PDF HTML 收藏
2604.23542 2026-04-28 cs.CV

AusSmoke meets MultiNatSmoke: a fully-labelled diverse smoke segmentation dataset

AusSmoke 遇见 MultiNatSmoke:一个完全标注的多样化烟雾分割数据集

Weihao Li, Hongjin Zhao, Gao Zhu, Ge-Peng Ji, Nicholas Wilson, Marta Yebra, Nick Barnes

机构 * Bushfire Research Centre of Excellence(火灾研究卓越中心)

AI总结 本文提出AusSmoke和MultiNatSmoke数据集,通过整合国际公开数据和澳大利亚实地采集数据,解决烟雾分割数据集规模小、地域局限和依赖合成图像的问题,并验证了模型在不同地理环境下的性能提升。

Comments Accepted to WACV 2026. Project page: https://github.com/henryzhao0615/MultiNatSmoke

URL PDF HTML 收藏
2509.18831 2026-04-22 cs.GR cs.AI cs.CV cs.LG cs.MM

Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters

Text Slider: 一种高效且即插即用的连续概念控制框架,通过LoRA适配器实现图像/视频合成

Pin-Yen Chiu, I-Sheng Fang, Jun-Cheng Chen

机构 * Research Center for Information Technology Innovation, Academia Sinica(学术院信息科技创新研究中心)

AI总结 Text Slider通过LoRA适配器在预训练文本编码器中识别低秩方向,实现图像和视频合成中的连续概念控制,提升效率并减少训练时间和GPU内存消耗。

Comments Accepted by WACV 2026. We provide more experimental results on the train-free version of our algorithm. Project page: https://textslider.github.io/ Code: https://github.com/aiiu-lab/TextSlider

URL PDF HTML 收藏