arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

共收录 2109
2512.15949 2025-12-19 cs.CV

The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs

感知观测站:多模态大语言模型的鲁棒性与基础性表征

Tejas Anvekar, Fenil Bardoliya, Pavan K. Turaga, Chitta Baral, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学)

AI总结 感知观测站通过系统性扰动和真实数据集,评估多模态大语言模型在视觉基础性和关系结构上的鲁棒性,揭示其在扰动下的表现,为模型分析提供系统基础。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15707 2025-12-18 cs.CV

GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection

GateFusion:用于活动说话检测的分层门控跨模态融合

Yu Wang, Juhyung Ha, Frangil M. Ramirez, Yuchen Wang, David J. Crandall

机构 * Indiana University(印第安纳大学)

AI总结 GateFusion通过分层门控融合解码器提升活动说话检测的跨模态融合效果,实现新的SOTA结果。

Comments accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15581 2025-12-18 cs.CV cs.LG

IMKD: Intensity-Aware Multi-Level Knowledge Distillation for Camera-Radar Fusion

IMKD:基于多级知识蒸馏的强度感知相机-雷达融合

Shashank Mishra, Karan Patil, Didier Stricker, Jason Rambach

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

AI总结 IMKD通过多级知识蒸馏提升雷达-相机融合性能,实现67.0% NDS和61.0% mAP的高精度检测

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026. 22 pages, 8 figures. Includes supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15006 2025-12-18 cs.CV cs.AI

Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation

评估视频问题生成在专家知识获取中的能力

Huaying Zhang, Atsushi Hashimoto, Tosho Hirasawa

机构 * OMRON SINIC X Corp.(OMRON SINIC X公司) Hokkaido University(北海道大学)

AI总结 本文提出了一种评估视频问题生成能力的新协议,通过模拟专家问答交流来评估问题质量,利用新构建的EgoExoAsk数据集验证了模型在获取未见知识方面的有效性。

Comments WACV 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14994 2025-12-18 cs.CV cs.AI

Where is the Watermark? Interpretable Watermark Detection at the Block Level

水印在哪里?基于块级的可解释水印检测

Maria Bulychev, Neil G. Marchant, Benjamin I. P. Rubinstein

机构 * University of Melbourne(墨尔本大学)

AI总结 本文提出一种基于块级的可解释水印检测方法,通过局部嵌入与区域可解释性结合,在保持鲁棒性的同时提供更清晰的检测结果。

Comments 20 pages, 14 figures. Camera-ready for WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07167 2025-12-18 cs.CV

Cascaded Dual Vision Transformer for Accurate Facial Landmark Detection

级联双视觉变换器用于精确面部 landmark 检测

Ziqiang Dang, Jianfang Li, Lin Liu

机构 * Zhejiang University(浙江大学) Institute for Intelligent Computing(智能计算研究院) Alibaba Group(阿里巴巴集团)

AI总结 本文提出级联双视觉变换器,通过通道分割和长跳连接提升面部 landmark 检测精度,优于现有最佳方法。

Comments Accepted by WACV 2025. The code can be found at https://github.com/Human3DAIGC/AccurateFacialLandmarkDetection . Supplementary material is included at the end of the main paper (3 pages, 5 figures, 2 tables)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19534 2025-12-16 cs.CV cs.LG

QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain

QUOTA: 通过文本到图像模型对任意领域的对象进行量化

Wenfang Sun, Yingjun Du, Gaowen Liu, Yefeng Zheng, Cees G. M. Snoek

机构 * University of Amsterdam(阿姆斯特丹大学) Cisco Research(思科研究) Westlake University(西湖大学)

AI总结 QUOTA通过双循环元学习策略,实现无需重新训练即可在任意领域中高效量化对象数量。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12963 2025-12-16 cs.CV

SCAdapter: Content-Style Disentanglement for Diffusion Style Transfer

SCAdapter: 内容-风格解耦用于扩散风格迁移

Luan Thanh Trinh, Kenji Doi, Atsuki Osanai

机构 * LY Corporation(LY公司)

AI总结 SCAdapter通过CLIP图像空间实现内容与风格的解耦,提升扩散模型在逼真图像迁移中的效果和效率。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12610 2025-12-16 cs.CV cs.IR

Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching

基于补丁的检索:实例级匹配的实用技术集合

Wonseok Choi, Sohwi Lim, Nam Hyeon-Woo, Moon Ye-Bin, Dong-Ju Jeong, Jinyoung Hwang, Tae-Hyun Oh

机构 * POSTECH KAIST(韩国科学技术院) Samsung Research(三星研究所)

AI总结 Patchify通过基于补丁的检索框架实现高效实例级匹配,结合LocScore评估空间正确性,提升检索性能与可解释性。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12296 2025-12-16 cs.CV cs.LG

GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search

GrowTAS: 从小到大逐步扩展子网络以实现高效的ViT架构搜索

Hyunju Lee, Youngmin Oh, Jeimin Jeon, Donghyeon Baek, Bumsub Ham

机构 * Yonsei University(延世大学)

AI总结 GrowTAS通过逐步扩展子网络,提高ViT架构搜索的效率和稳定性。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11901 2025-12-16 cs.CV cs.LG

CLARGA: Multimodal Graph Representation Learning over Arbitrary Sets of Modalities

CLARGA:任意模态集合上的多模态图表示学习

Santosh Patapati

机构 * Santosh Patapati(独立研究者)

AI总结 CLARGA是一种通用的多模态融合架构,通过构建注意力加权图实现多模态表示学习,适用于任意模态集合,具有高效的融合能力和良好的鲁棒性。

Comments WACV; Supplementary material is available on CVF proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11894 2025-12-16 cs.CV cs.LG

mmWEAVER: Environment-Specific mmWave Signal Synthesis from a Photo and Activity Description

mmWEAVER: 从照片和活动描述合成环境特定的毫米波信号

Mahathir Monjur, Shahriar Nirjon

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校)

AI总结 mmWeaver通过隐式神经表示和超网络,高效生成环境特定的毫米波信号,提升活动识别和姿态估计性能。

Comments Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026 (WACV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07935 2025-12-16 cs.CV cs.AI

DiffRegCD: Integrated Registration and Change Detection with Diffusion Features

DiffRegCD:集成注册与变化检测的扩散特征

Seyedehanita Madani, Rama Chellappa, Vishal M. Patel

机构 * Johns Hopkins University(约翰霍普金斯大学)

AI总结 DiffRegCD通过结合扩散特征与分类任务,实现统一的注册与变化检测,提升在复杂场景下的鲁棒性和精度。

Comments 10 pages, 6 figures. Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11719 2025-12-15 cs.CV

Referring Change Detection in Remote Sensing Imagery

遥感图像中的指称变化检测

Yilmaz Korkmaz, Jay N. Paranjape, Celso M. de Melo, Vishal M. Patel

机构 * Johns Hopkins University(约翰霍普金斯大学) DEVCOM U.S. Army Research Laboratory(美国陆军研究实验室)

AI总结 本文提出指称变化检测方法,通过自然语言提示实现遥感图像中特定类别的变化检测,并引入两阶段框架提升数据生成效率。

Comments 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11215 2025-12-15 cs.CV

SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection

SmokeBench: 评估多模态大语言模型用于野火烟雾检测

Tianye Qi, Weihao Li, Nick Barnes

机构 * Australian National University(澳大利亚国立大学)

AI总结 SmokeBench评估多模态大语言模型在野火烟雾检测中的性能,发现模型在烟雾定位方面存在显著局限,尤其在早期阶段。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07299 2025-12-15 cs.CV

VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models

VADER:基于关系感知大型语言模型的因果视频异常理解

Ying Cheng, Yu-Ho Lin, Min-Hung Chen, Fu-En Yang, Shang-Hong Lai

机构 * National Tsing Hua University(清华大学) NVIDIA

AI总结 VADER通过整合关系感知的大型语言模型与视频关键帧特征,提升视频异常的因果理解与解释能力。

Comments Accepted to WACV 2026. Project page available at: https://vader-vau.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16713 2025-12-15 cs.CV

Conditional Text-to-Image Generation with Reference Guidance

基于参考引导的条件文本到图像生成

Taewook Kim, Ze Wang, Zhengyuan Yang, Jiang Wang, Lijuan Wang, Zicheng Liu, Qiang Qiu

机构 * Purdue University(普渡大学) AMD Microsoft(微软)

AI总结 本文提出基于参考引导的条件文本到图像生成方法,通过专家插件提升模型在文本拼写和多语言生成等任务上的表现。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10939 2025-12-12 cs.CV

GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting

GaussianHeadTalk: 无抖动的3D说话头:基于音频驱动的高斯点云法

Madhav Agarwal, Mingtian Zhang, Laura Sevilla-Lara, Steven McDonagh

机构 * University of Edinburgh(爱丁堡大学) University College London(伦敦大学学院)

AI总结 本文提出基于音频驱动的3D说话头生成方法,利用高斯点云法和Transformer模型实现高保真、实时的虚拟角色生成。

Comments IEEE/CVF Winter Conference on Applications of Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10321 2025-12-12 cs.CV

Point2Pose: A Generative Framework for 3D Human Pose Estimation with Multi-View Point Cloud Dataset

Point2Pose:基于多视角点云数据集的3D人体姿态估计生成框架

Hyunsoo Lee, Daeum Jeon, Hyeokjae Oh

机构 * ECE, Seoul National University(首尔国立大学电子与计算机工程系) CS, KAIST(韩国科学技术院计算机科学系) Soulart Inc.(Soulart公司)

AI总结 Point2Pose通过生成模型和大规模数据集提升3D人体姿态估计的准确性与鲁棒性。

Comments WACV 2026 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09383 2025-12-12 cs.CV

Perception-Inspired Color Space Design for Photo White Balance Editing

受感知启发的色彩空间设计用于照片白平衡编辑

Yang Cheng, Ziteng Cui, Shenghan Su, Lin Gu, Zenghui Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) The University of Tokyo(东京大学) Tohoku University(东北大学)

AI总结 本文提出了一种基于受感知启发的可学习HSI色彩空间的白平衡校正框架,以解决传统加色模型在复杂光照条件下的局限性。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07730 2025-12-12 cs.CV cs.AI

SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination

SAVE:基于稀疏自编码器的视觉信息增强用于缓解物体幻觉

Sangha Park, Seungryong Yoo, Jisoo Mok, Sungroh Yoon

机构 * Department of Electrical and Computer Engineering, Seoul National University(电子与计算机工程系,首尔国立大学) Daegu Gyeongbuk Institute of Science and Technology(大邱庆州科学技术院) IPAI, AIIS, ASRI, INMC, and ISRC, Seoul National University(IPAI、AIIS、ASRI、INMC 和 ISRC,首尔国立大学)

AI总结 SAVE通过引导模型沿稀疏自编码器潜在特征减少物体幻觉,提升视觉理解能力。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07702 2025-12-12 cs.CV cs.AI

Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment

引导不应生成的内容:用于文本-图像对齐的自动化负提示

Sangha Park, Eunji Kim, Yeongtak Oh, Jooyoung Choi, Sungroh Yoon

机构 * Department of Electrical and Computer Engineering, Seoul National University(电子与计算机工程系,首尔国立大学) Amazon(亚马逊) IPAI, AIIS, ASRI, INMC, and ISRC, Seoul National University(IPAI、AIIS、ASRI、INMC 和 ISRC,首尔国立大学)

AI总结 本文提出NPC方法,通过自动化负提示生成提升文本-图像对齐效果,在GenEval++和Imagine-Bench上取得优异成绩。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00846 2025-12-12 cs.CV

AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent

AFRAgent:一种基于自适应特征归一化的高分辨率感知GUI代理

Neeraj Anand, Rishabh Jain, Sohan Patnaik, Balaji Krishnamurthy, Mausoom Sarkar

机构 * Media and Data Science Research, Adobe(Adobe媒体与数据科学研究所)

AI总结 AFRAgent通过自适应特征归一化技术,在保持高分辨率细节的同时,实现了更高效的GUI自动化性能,比现有模型小四分之一。

Comments Accepted at WACV 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00087 2025-12-12 cs.CV

Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data

探索多模态课堂数据中教学活动和话语的自动化识别

Ivo Bueno, Ruikun Hou, Babette Bühler, Tim Fütterer, James Drimalla, Jonathan Kyle Foster, Peter Youngs, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci

AI总结 本文通过多模态分析方法,实现了课堂活动中教学活动和话语的自动化识别,展示了微调模型在视频和 transcripts 上的高准确率,为可扩展的教师反馈系统提供了基础。

Comments This article has been accepted for publication in the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21888 2025-12-12 cs.CV

CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding

CAPE:一种基于CLIP的互补热图线索点集用于具身参照理解

Fevziye Irem Eyiokur, Dogucan Yaman, Hazım Kemal Ekenel, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Istanbul Technical University(伊斯坦布尔技术大学) Carnegie Mellon University(卡内基梅隆大学) KIT Campus Transfer GmbH (KCT)(KIT校园转移有限责任公司)

AI总结 CAPE通过双模型框架和CLIP-aware Pointing Ensemble模块,提升具身参照理解任务中指向线索的多模态推理能力,实现75.0 mAP的高精度表现。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13901 2025-12-12 cs.CV

Dressing the Imagination: A Dataset for AI-Powered Translation of Text into Fashion Outfits and A Novel NeRA Adapter for Enhanced Feature Adaptation

为想象力着装:一个用于人工智能驱动的文本到时尚装扮的数据库及一种新的NeRA适配器用于增强特征适应

Gayatri Deshmukh, Somsubhra De, Chirag Sehgal, Jishu Sen Gupta, Sparsh Mittal

机构 * IIT Madras(印度理工学院Madras分校) Delhi Technological University(德里技术大学) IIT BHU(印度理工学院BHU分校) IIT Roorkee(印度理工学院Roorkee分校)

AI总结 FLORA数据集和NeRA适配器旨在提升人工智能生成时尚设计的精度与风格丰富度。

Comments Accepted as a Conference Paper at WACV 2026 (USA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10102 2025-12-12 cs.CV

Hierarchical Instance Tracking to Balance Privacy Preservation with Accessible Information

层级实例跟踪以平衡隐私保护与可获取信息

Neelima Prasad, Jarek Reynolds, Neel Karsanbhai, Tanusree Sharma, Lotus Zhang, Abigale Stangl, Yang Wang, Leah Findlater, Danna Gurari

机构 * University of Colorado Boulder(科罗拉多大学博尔德分校) Pennsylvania State University(宾夕法尼亚州立大学) University of Washington(华盛顿大学) Georgia Institute of Technology(佐治亚理工学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文提出层级实例跟踪任务,构建首个支持该任务的基准数据集,通过评估多种模型展示数据集的挑战性,旨在平衡隐私保护与信息可获取性。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09847 2025-12-11 cs.CV

From Detection to Anticipation: Online Understanding of Struggles across Various Tasks and Activities

从检测到预判:跨多种任务和活动的在线理解挣扎

Shijia Feng, Michael Wray, Walterio Mayol-Cuevas

机构 * University of Bristol(布里斯托大学) Amazon(亚马逊)

AI总结 本文提出了一种在线检测和预判挣扎的方法,通过改进模型实现跨任务和活动的实时应用,提升辅助系统的响应能力。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09792 2025-12-11 cs.CV

FastPose-ViT: A Vision Transformer for Real-Time Spacecraft Pose Estimation

FastPose-ViT:一种用于实时航天器姿态估计的视觉Transformer

Pierre Ancey, Andrew Price, Saqib Javed, Mathieu Salzmann

机构 * EPFL(瑞士联邦理工学院) Swiss Data Science Center(瑞士数据科学中心)

AI总结 FastPose-ViT提出了一种基于视觉Transformer的实时航天器姿态估计方法,通过直接回归6DoF姿态,实现了高效且高精度的姿态估计。

Comments Accepted to WACV 2026. Preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09095 2025-12-11 cs.CV

Food Image Generation on Multi-Noun Categories

多名词类别的食物图像生成

Xinyue Pan, Yuhao Chen, Jiangpeng He, Fengqing Zhu

机构 * Purdue University(普渡大学) University of Waterloo(滑铁卢大学)

AI总结 本文提出FoCULR方法,通过整合食品领域知识和早期引入核心概念,解决多名词类别食物图像生成中的语义误解和布局错误问题。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏