arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

共收录 2109
2507.18594 2025-12-09 cs.CV cs.AI cs.LG

DRWKV: Focusing on Object Edges for Low-Light Image Enhancement

DRWKV:聚焦于物体边缘的低光照图像增强

Xuecheng Bai, Yuxiang Wang, Boyu Hu, Qinyuan Jie, Chuanzhi Xu, Kechen Li, Hongru Xiao, Vera Chung

机构 * Shenyang Ligong University(沈阳理工大学) The University of Sydney(悉尼大学) University of International Business and Economics(国际商务经济大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Tongji University(同济大学)

AI总结 DRWKV通过整合全局边缘视网膜理论和进化WKV注意力机制,有效提升低光照图像增强的边缘保真度和视觉自然度,实现高PSNR、SSIM和NIQE性能,并在多目标跟踪任务中展现良好的泛化能力。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14670 2025-12-09 cs.CV

Gene-DML: Dual-Pathway Multi-Level Discrimination for Gene Expression Prediction from Histopathology Images

Gene-DML:双路径多级判别用于从组织病理图像预测基因表达

Yaxuan Song, Jianan Fan, Hang Chang, Weidong Cai

机构 * The University of Sydney, Australia(悉尼大学) Lawrence Berkeley National Laboratory, USA(伯克利国家实验室)

AI总结 Gene-DML通过双路径多级判别方法,提升组织病理图像与基因表达谱的跨模态对齐,实现高精度的基因表达预测。

Comments Accepted by The IEEE/CVF Winter Conference on Applications of Computer Vision 2026 (WACV2026). Code and data available at https://github.com/YXSong000/Gene-DML

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05762 2025-12-08 cs.CV cs.GR

FNOPT: Resolution-Agnostic, Self-Supervised Cloth Simulation using Meta-Optimization with Fourier Neural Operators

FNOPT:基于元优化的四ier神经算子的自监督布料模拟

Ruochen Chen, Thuy Tran, Shaifali Parashar

机构 * CNRS, École Centrale de Lyon, INSA Lyon, Université Claude Bernard Lyon 1, LIRIS, UMR5205, France(法国国家科学研究中心、里昂中央理工大学、里昂国家应用科学学院、里昂第一大学、LIRIS、UMR5205)

AI总结 FNOpt通过元优化和傅里叶神经算子实现自监督布料模拟,无需重新训练即可在不同分辨率和运动模式下保持稳定和准确。

Comments Accepted for WACV

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05140 2025-12-08 cs.CV cs.AI

FlowEO: Generative Unsupervised Domain Adaptation for Earth Observation

FlowEO: 生成式无监督领域适应用于地球观测

Georges Le Bellier, Nicolas Audebert

AI总结 FlowEO通过生成模型实现地球观测图像的无监督领域适应,有效应对多源异构数据的适应挑战,提升遥感图像分类和分割性能。

Comments 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Mar 2026, Tucson (AZ), United States

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01498 2025-12-08 cs.CV cs.AI

AortaDiff: A Unified Multitask Diffusion Framework For Contrast-Free AAA Imaging

AortaDiff:一种用于无对比剂AAA成像的统一多任务扩散框架

Yuxuan Ou, Ning Bi, Jiazhen Pan, Jiancheng Yang, Boliang Yu, Usama Zidan, Regent Lee, Vicente Grau

机构 * Department of Engineering Science, University of Oxford, United Kingdom(牛津大学工程科学系) Nuffield Department of Surgical Sciences, University of Oxford, United Kingdom(牛津大学外科科学努尔菲尔德系) Technical University of Munich, Germany(慕尼黑技术大学) ELLIS Institute Finland, Finland(芬兰ELLIS研究所) Aalto University, Finland(阿alto大学)

AI总结 AortaDiff通过统一多任务扩散框架实现无对比剂AAA成像,同时生成合成CECT图像并分割主动脉腔和血栓,提升了分割精度和临床测量准确性。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13275 2025-12-08 cs.CV

ChartQA-X: Generating Explanations for Visual Chart Reasoning

ChartQA-X: 生成视觉图表推理的解释

Shamanthak Hegde, Pooyan Fazli, Hasti Seifi

机构 * Arizona State University(亚利桑那州立大学)

AI总结 ChartQA-X通过生成图表解释提升视觉推理,其数据集包含30,799个图表样本,模型在解释质量和问答准确率上表现优异。

Comments WACV 2026. Project Page: https://teal-lab.github.io/chartqa-x

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04504 2025-12-08 cs.CV

AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM

AnyAnomaly: 零样本可定制化视频异常检测与LVLM

Sunghyun Ahn, Youngwan Jo, Kijung Lee, Sein Kwon, Inpyo Hong, Sanghyun Park

机构 * Yonsei University(延世大学)

AI总结 AnyAnomaly通过上下文感知的视觉问答模型实现零样本可定制化视频异常检测,无需微调大型视觉语言模型,在多个基准测试中取得最佳性能。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04927 2025-12-05 cs.CV

Virtually Unrolling the Herculaneum Papyri by Diffeomorphic Spiral Fitting

通过仿射螺旋拟合虚拟展开赫库拉尼姆莎草纸

Paul Henderson

机构 * University of Glasgow(格拉斯哥大学)

AI总结 通过仿射螺旋拟合方法实现莎草纸的虚拟展开,自动拟合表面模型以生成连续的2D展开表示。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07828 2025-12-05 cs.CV

MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions

MMHOI:建模复杂3D多人类多物体交互

Kaen Kogashi, Anoop Cherian, Meng-Yu Jennifer Kuo

机构 * Mitsubishi Electric Japan(三菱电机日本公司) Mitsubishi Electric Research Labs(三菱电机研究实验室) Nara Women’s University(奈良女子大学)

AI总结 MMHOI提出了一种大规模多人类多物体交互数据集和端到端Transformer网络,用于建模复杂3D人类-物体交互,实现了最先进的性能。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04487 2025-12-05 cs.CV

Controllable Long-term Motion Generation with Extended Joint Targets

通过扩展目标关节实现可控的长期运动生成

Eunjong Lee, Eunhee Kim, Sanghoon Hong, Eunho Jung, Jihoon Kim

机构 * Cinamon Inc.(Cinamon公司)

AI总结 COMET通过扩展目标关节实现可控的长期运动生成,利用高效Transformer基于条件VAE实现精确交互控制,并通过参考引导反馈机制确保长期稳定性。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04356 2025-12-05 cs.CV cs.AI cs.CL cs.LG

Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment

通过自增强对比对齐缓解多模态大语言模型中的对象和动作幻觉

Kai-Po Chang, Wei-Yuan Cheng, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国家交通大学通信工程研究所) NVIDIA

AI总结 SANTA框架通过自增强对比对齐方法,有效缓解多模态大语言模型中的对象和动作幻觉问题。

Comments IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026. Project page: https://kpc0810.github.io/santa/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19717 2025-12-05 cs.CV cs.AI cs.LG cs.RO

MonoPP: Metric-Scaled Self-Supervised Monocular Depth Estimation by Planar-Parallax Geometry in Automotive Applications

MonoPP: 基于平面-视差几何的汽车应用中利用度量尺度自监督单目深度估计

Gasser Elazab, Torben Gräber, Michael Unterreiner, Olaf Hellwich

机构 * CARIAD SE(CARIAD公司) Technische Universität Berlin(柏林技术大学)

AI总结 MonoPP通过平面-视差几何实现自监督单目深度估计,利用多帧和单帧网络及姿态网络,实现汽车应用中度量尺度深度预测的先进性能。

Comments Accepted at WACV 25, project page: https://mono-pp.github.io/

Journal ref Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Tucson, AZ, USA, 26 February 2025, pp. 2777-2787

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03749 2025-12-04 cs.CV

Fully Unsupervised Self-debiasing of Text-to-Image Diffusion Models

文本到图像扩散模型的完全无监督自去偏

Korada Sri Vardhana, Shrikrishna Lolla, Soma Biswas

机构 * Indian Institute of Science(印度科学研究院)

AI总结 SelfDebias是一种完全无监督的文本到图像扩散模型去偏方法,通过识别语义聚类来减少生成图像中的偏见,同时保持视觉真实性。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03317 2025-12-04 cs.CV cs.AI cs.LG cs.RO

NavMapFusion: Diffusion-based Fusion of Navigation Maps for Online Vectorized HD Map Construction

NavMapFusion: 基于扩散的导航地图融合用于在线向量化的高精度地图构建

Thomas Monninger, Zihan Zhang, Steffen Staab, Sihao Ding

机构 * Mercedes-Benz Research & Development North America, USA(梅赛德斯-奔驰北美研究与开发) University of Stuttgart, Germany(斯图加特大学) University of California, San Diego, USA(加州大学圣地亚哥分校) University of Southampton, United Kingdom(南安普顿大学)

AI总结 NavMapFusion通过结合低保真先验与高保真传感器数据,利用扩散模型实现在线高精度地图构建,提升环境表示的准确性和实时性。

Comments Accepted to 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13430 2025-12-04 cs.CV cs.AI cs.LG cs.RO

AugMapNet: Improving Spatial Latent Structure via BEV Grid Augmentation for Enhanced Vectorized Online HD Map Construction

AugMapNet: 通过BEV网格增强改进空间潜在结构以提升矢量化的在线HD地图构建

Thomas Monninger, Md Zafar Anwar, Stanislaw Antol, Steffen Staab, Sihao Ding

机构 * Mercedes-Benz Research & Development North America, USA(梅赛德斯-奔驰北美研究与开发)

AI总结 AugMapNet通过BEV网格增强改进空间潜在结构,提升矢量化在线HD地图构建的性能和结构化表示。

Comments Accepted to 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02952 2025-12-03 cs.CV

Layout Anything: One Transformer for Universal Room Layout Estimation

布局万物:一种用于通用房间布局估计的变换器

Md Sohag Mia, Muhammad Abdullah Adnan

机构 * Nanjing University of Information Science and Technology(南京信息工程大学) Bangladesh University of Engineering and Technology(孟加拉国工程与技术大学)

AI总结 Layout Anything通过整合任务条件查询和对比学习,提出了一种基于变换器的通用房间布局估计框架,实现了高速推理和高精度性能。

Comments Published at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02846 2025-12-03 cs.CV

Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?

瞬间动作预见:多模态线索能替代视频到何种程度?

Manuel Benavent-Lledo, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez

机构 * Universidad de Alicante(阿利坎特大学) Foundation for Research and Technology-Hellas(希腊基础研究与技术基金会) University of Crete(克里特大学) Hellenic Mediterranean University(希腊地中海大学)

AI总结 AAG通过结合单帧RGB特征与深度线索及先前动作信息,实现了多模态单帧动作预见,能与视频聚合基线和先进方法在教学活动数据集上竞争。

Comments Accepted in WACV 2026 - Applications Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02737 2025-12-03 cs.CV cs.LG

Beyond Paired Data: Self-Supervised UAV Geo-Localization from Reference Imagery Alone

超越配对数据:仅从参考影像进行无人机地理定位的自监督方法

Tristan Amadei, Enric Meinhardt-Llopis, Benedicte Bascle, Corentin Abgrall, Gabriele Facciolo

机构 * Thales LAS Universite Paris-Saclay(巴黎萨克雷大学) ENS Paris-Saclay(巴黎萨克雷高等学院) CNRS(国家科学研究中心) Centre Borelli(Borelli中心)

AI总结 本文提出了一种无需配对数据的自监督无人机地理定位方法,通过卫星参考影像训练,有效提升了在GNSS拒止环境下的定位性能。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02727 2025-12-03 cs.CV cs.AI cs.LG

DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions

DF-Mamba:用于手姿估计的可变形状态空间建模

Yifan Zhou, Takehiko Ohkawa, Guwenxiao Zhou, Kanoko Goto, Takumi Hirose, Yusuke Sekikawa, Nakamasa Inoue

机构 * Institute of Science Tokyo(东京科学研究所) Denso IT Laboratory(Denso IT实验室) The University of Tokyo(东京大学)

AI总结 DF-Mamba通过可变形状态空间建模提升3D手姿估计的鲁棒性和精度,优于现有方法。

Comments Accepted to WACV 2026. Project page: https://tkhkaeio.github.io/projects/25-dfmamba/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08294 2025-12-03 cs.CV

SkelSplat: Robust Multi-view 3D Human Pose Estimation with Differentiable Gaussian Rendering

SkelSplat: 基于可微高斯渲染的鲁棒多视角3D人体姿态估计

Laura Bragagnolo, Leonardo Barcellona, Stefano Ghidoni

机构 * University of Padova(帕多瓦大学) University of Amsterdam(阿姆斯特丹大学)

AI总结 SkelSplat通过可微高斯渲染实现多视角3D人体姿态估计,无需3D真实数据监督,有效提升跨数据集泛化能力与遮挡鲁棒性。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02438 2025-12-03 cs.CV cs.AI

Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources

通过有限计算资源下的动量自蒸馏提升医学视觉-语言预训练

Phuc Pham, Nhu Pham, Ngoc Quoc Ly

机构 * Faculty of Information Technology, University of Science, Ho Chi Minh City, Vietnam(信息科技学院,科学大学,胡志明市,越南) Vietnam National University, Ho Chi Minh City, Vietnam(越南国家大学,胡志明市,越南)

AI总结 本研究通过动量自蒸馏与梯度累积提升医学视觉-语言预训练效率,实现高效多模态学习并提升零样本分类和少量样本适应性能。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13922 2025-12-03 eess.IV cs.CV cs.LG cs.MM

Self-Supervised Compression and Artifact Correction for Streaming Underwater Imaging Sonar

面向流式水下成像声纳的自监督压缩与伪影校正

Rongsheng Qian, Chi Xu, Xiaoqiang Ma, Hao Fang, Yili Jin, William I. Atlas, Jiangchuan Liu

机构 * Simon Fraser University(西蒙弗雷泽大学) Douglas College(道格拉斯学院) McGill University(麦吉尔大学) Wild Salmon Center(野生鲑鱼中心)

AI总结 SCOPE通过自监督方法实现低比特率水下声纳流式传输,有效压缩并校正声纳伪影,提升实时检测性能。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19478 2025-12-03 cs.CV

Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment

通过无监督帧到段对齐实现感知-aware的动作分割

Quoc-Huy Tran, Ahmed Mehmood, Muhammad Ahmed, Muhammad Naufil, Anas Zafar, Andrey Konin, M. Zeeshan Zia

机构 * Retrocausal, Inc.(Retrocausal公司)

AI总结 本文提出一种无监督框架,通过帧到段对齐实现感知-aware的动作分割,利用变换器模型和伪标签提升分割性能。

Comments Accepted to WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11508 2025-12-02 cs.CV

Towards Fast and Scalable Normal Integration using Continuous Components

通过连续组件实现快速且可扩展的法线积分

Francesco Milano, Jen Jen Chung, Lionel Ott, Roland Siegwart

机构 * ETH Zurich(苏黎世联邦理工学院) The University of Queensland(昆士兰大学)

AI总结 本文提出通过连续组件估计实现快速且可扩展的法线积分,显著提升处理大分辨率法线图的效率。

Comments Accepted by the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026, first round. Camera-ready version. 17 pages, 9 figures, 6 tables. Code is available at https://github.com/francescomilano172/normal_integration_continuous_components

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12168 2025-12-02 cs.CV cs.GR

Sketch-guided Cage-based 3D Gaussian Splatting Deformation

基于草图的基于笼子的3D高斯点扩散变形

Tianhao Xie, Noam Aigerman, Eugene Belilovsky, Tiberiu Popa

机构 * Concordia University(康科德大学) Université de Montréal(蒙特利尔大学) Mila(Mila研究院)

AI总结 本文提出了一种基于草图的3D GS变形系统,结合笼子变形与神经雅可比场,实现对3D模型几何的精细控制,并通过实验展示其在静态模型动画化中的应用。

Comments 10 pages, 9 figures, accepted at WACV 26, project page: https://tianhaoxie.github.io/project/gs_deform/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00936 2025-12-02 cs.CV

SceneProp: Combining Neural Network and Markov Random Field for Scene-Graph Grounding

SceneProp:结合神经网络和马尔可夫随机场进行场景图接地

Keita Otani, Tatsuya Harada

机构 * The University of Tokyo(东京大学) RIKEN AIP(日本科学技术研究所(RIKEN)先进研究所(AIP))

AI总结 SceneProp通过将场景图接地重新公式化为马尔可夫随机场的MAP推断问题,有效提升了复杂查询下的视觉接地性能。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00909 2025-12-02 cs.CV

TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model

TalkingPose: 基于反馈引导扩散模型的高效人脸与手势动画

Alireza Javanmardi, Pragati Jaiswal, Tewodros Amberbir Habtegebrial, Christen Millerdurai, Shaoxiang Wang, Alain Pagani, Didier Stricker

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) RPTU(鲁尔大学)

AI总结 TalkingPose通过反馈引导扩散模型实现高效的人体上半身动画生成,支持无限持续时间的连贯动画创作。

Comments WACV 2026, Project page available at https://dfki-av.github.io/TalkingPose

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10799 2025-12-02 cs.CV

GFT: Graph Feature Tuning for Efficient Point Cloud Analysis

GFT:用于高效点云分析的图特征调谐

Manish Dhakal, Venkat R. Dasari, Rajshekhar Sunderraman, Yi Ding

机构 * Department of Computer Science, Georgia State University(计算机科学系,佐治亚州立大学) DEVCOM Army Research Laboratory(陆军研究实验室)

AI总结 GFT通过图特征调谐方法,针对点云数据高效减少可训练参数,提升点云分析任务的效率与性能。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08819 2025-12-02 eess.IV cs.CV

Boosting Diffusion Guidance via Learning Degradation-Aware Models for Blind Super Resolution

通过学习降质感知模型提升扩散引导:用于盲超分辨率

Shao-Hao Lu, Ren Wang, Ching-Chun Huang, Wei-Chen Chiu

AI总结 本文提出DADiff,通过学习降质感知模型提升扩散引导,解决盲超分辨率中的降质核依赖问题,实现高保真结果。

Comments To appear in WACV 2025. Code is available at: https://github.com/ryanlu2240/DADiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14672 2025-12-02 cs.CV

Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology

通过形态学优化对抗数据中的不可行包含以实现语义分割

Shamik Basu, Luc Van Gool, Christos Sakaridis

机构 * University Of Bologna(博洛尼亚大学) INSAIT(INSAIT研究所) Sofia University St. Kliment Ohridski(索菲亚大学圣克莱门特·欧赫里迪斯基大学) ETH Zürich(苏黎世联邦理工学院)

AI总结 InSeIn通过提取空间类关系约束并采用可微形态学损失优化语义分割,提升预测可行性与性能。

Comments Published in 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏