arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

至 收录 2109
2409.20469 2026-02-26 cs.CV

PoseAdapt: Sustainable Human Pose Estimation via Continual Learning Benchmarks and Toolkit

PoseAdapt: 通过持续学习基准和工具包实现可持续的人体姿态估计

Muhammad Saif Ullah Khan, Didier Stricker

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

AI总结 PoseAdapt通过持续学习基准和工具包实现可持续的人体姿态估计,提供持续学习方法的评估和模型适应框架,以提高模型适应性并减少重复训练需求。

Comments Accepted in WACV 2026 Applications Track

URL PDF HTML 收藏
2602.19706 2026-02-24 cs.CV

HDR Reconstruction Boosting with Training-Free and Exposure-Consistent Diffusion

通过无训练和曝光一致的扩散实现HDR重建增强

Yo-Tin Lin, Su-Kai Chen, Hou-Ning Hu, Yen-Yu Lin, Yu-Lun Liu

机构 * National Yang Ming Chiao Tung University(国家阳明交通大学) MediaTek Inc.(联发科公司)

AI总结 通过无训练和曝光一致的扩散技术提升HDR重建效果,增强过曝区域的图像质量并保持多曝光一致性。

Comments WACV 2026. Project page: https://github.com/EusdenLin/HDR-Reconstruction-Boosting

URL PDF HTML 收藏
2511.06450 2026-02-24 cs.CV cs.LG

Countering Multi-modal Representation Collapse through Rank-targeted Fusion

通过秩目标融合对抗多模态表示崩溃

Seulgi Kim, Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 本文提出Rank-enhancing Token Fuser框架,通过提升有效秩对抗多模态表示崩溃,验证了深度与RGB融合的平衡性,并在动作预测任务中取得显著性能提升。

Comments Accepted in 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

URL PDF HTML 收藏
2602.18720 2026-02-24 cs.CV

Subtle Motion Blur Detection and Segmentation from Static Image Artworks

从静态图像艺术品中检测和分割细微运动模糊

Ganesh Samarth, Sibendu Paul, Solale Tabarestani, Caren Chen

机构 * Amazon Prime Video(亚马逊Prime视频)

AI总结 本文提出SMBlurDetect,通过生成高质量运动模糊数据集和端到端检测器,实现多粒度零样本检测,提升静态图像中细微运动模糊的检测与分割性能。

Comments InProceedings of the Winter Conference on Applications of Computer Vision 2026

URL PDF HTML 收藏
2602.18618 2026-02-24 cs.CV

Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent Space

为你叙述:基于多纠缠潜在空间的提示引导音频视觉叙述面部生成

Aashish Chandra, Aashutosh A, Abhijit Das

机构 * Machine Intelligence Group, Department of CS&IS, BITS Pilani, Hyderabad Campus, India(机器智能组,计算机科学与信息学系,比斯汉学院海得拉巴校区,印度) Georgia Institute of Technology, USA(佐治亚理工学院,美国)

AI总结 本文提出了一种基于多纠缠潜在空间的音频视觉叙述面部生成方法,通过合成静态图像、语音档案和目标文本来生成逼真的说话面孔。

Comments To appear in the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026. Presented at Poster Session 1

URL PDF HTML 收藏
2602.14514 2026-02-23 cs.CV

Efficient Text-Guided Convolutional Adapter for the Diffusion Model

高效的文本引导卷积适配器用于扩散模型

Aryan Das, Koushik Biswas, Swalpa Kumar Roy, Badri Narayana Patro, Vinay Kumar Verma

机构 * VIT Bhopal(维特大学博帕尔分校) IIIT Delhi(德里印度理工学院) Tezpur University Assam(泰朱普大学阿萨姆分校) IIT Kanpur(坎普尔印度理工学院)

AI总结 本文提出Nexus Prime和Slim适配器,通过文本引导提升扩散模型的结构保持条件生成性能,显著减少参数量并保持高效果。

Comments Accepted in WACV 2026

URL PDF HTML 收藏
2602.14498 2026-02-23 cs.CV cs.LG

Uncertainty-Aware Vision-Language Segmentation for Medical Imaging

面向医学影像的不确定性感知视觉-语言分割

Aryan Das, Tanishq Rachamalla, Koushik Biswas, Swalpa Kumar Roy, Vinay Kumar Verma

机构 * VIT Bhopal(维特大学博帕尔分校) SAHE, Andhra Pradesh(安得拉邦SAHE) IIIT Delhi(德里印度理工学院) Tezpur University Assam(阿萨姆特兹普尔大学) IIT Kanpur(坎普尔印度理工学院)

AI总结 本文提出一种不确定性感知的多模态分割框架,通过引入SEU损失和MoDAB模块,提升医学影像分割的精度与效率。

Comments Accepted in WACV 2026

URL PDF HTML 收藏
2602.07835 2026-02-20 cs.CV

VFace: A Training-Free Approach for Diffusion-Based Video Face Swapping

VFace:一种无需训练的基于扩散模型的视频人脸交换方法

Sanoojan Baliah, Yohan Abeysinghe, Rusiru Thushara, Khan Muhammad, Abhinav Dhall, Karthik Nandakumar, Muhammad Haris Khan

AI总结 VFace提出一种无需训练的视频人脸交换方法,通过频率谱注意力插值、目标结构引导和流引导注意力时间平滑机制,提升时间一致性和视觉保真度。

Comments Accepted at WACV 2026

URL PDF HTML 收藏
2602.16669 2026-02-19 cs.CV

PredMapNet: Future and Historical Reasoning for Consistent Online HD Vectorized Map Construction

PredMapNet: 未来与历史推理用于一致的在线高精度向量地图构建

Bo Lang, Nirav Savaliya, Zhihao Zheng, Jinglun Feng, Zheng-Hang Yeh, Mooi Choo Chuah

机构 * Lehigh University(莱维大学) Honda Research Institute USA(本田美国研究院)

AI总结 PredMapNet通过结合地图实例跟踪和短期预测,实现一致的在线高精度向量地图构建,提升时间连续性和地图构建的稳定性。

Comments WACV 2026

URL PDF HTML 收藏
2602.16245 2026-02-19 cs.CV

HyPCA-Net: Advancing Multimodal Fusion in Medical Image Analysis

HyPCA-Net:在医学图像分析中推进多模态融合

J. Dhar, M. K. Pandey, D. Chakladar, M. Haghighat, A. Alavi, S. Mistry, N. Zaidi

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) RoentGen Health(RoentGen健康公司) Lulea University of Technology(卢勒奥大学) QUT(昆士兰科技大学) RMIT University(皇家墨尔本理工大学) Curtin University(Curtin大学) Deakin University(德肯大学)

AI总结 HyPCA-Net通过高效残差注意力模块和双视角级联注意力模块,提升多模态医学图像分析的性能与效率。

Comments Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026

URL PDF HTML 收藏
2602.16019 2026-02-19 cs.CV cs.AI

MedProbCLIP: Probabilistic Adaptation of Vision-Language Foundation Model for Reliable Radiograph-Report Retrieval

MedProbCLIP: 基于概率适应的视觉-语言基础模型用于可靠放射影像-报告检索

Ahmad Elallaf, Yu Zhang, Yuktha Priya Masupalli, Jeong Yang, Young Lee, Zechun Cao, Gongbo Liang

机构 * Texas A&M University-San Antonio(德克萨斯A&M大学-圣安东尼奥分校) Boise State University(博伊州立大学)

AI总结 MedProbCLIP 通过概率适应提升放射影像与报告检索的可靠性,优于现有基线方法。

Comments Accepted to the 2026 Winter Conference on Applications of Computer Vision (WACV) Workshops

URL PDF HTML 收藏
2602.15813 2026-02-18 cs.RO

FAST-EQA: Efficient Embodied Question Answering with Global and Local Region Relevancy

FAST-EQA: 有效体感问答与全局和局部区域相关性

Haochen Zhang, Nirav Savaliya, Faizan Siddiqui, Enna Sachdeva

机构 * Carnegie Mellon University(卡内基梅隆大学) Honda Research Institute USA(本田研究院美国)

AI总结 FAST-EQA通过结合全局和局部区域相关性,实现高效体感问答,提升回答可靠性并加快推理速度。

Comments WACV 2026

URL PDF HTML 收藏
2506.20367 2026-02-18 cs.GR cs.CV

DreamAnywhere: Object-Centric Panoramic 3D Scene Generation

DreamAnywhere: 基于对象的全景3D场景生成

Edoardo Alberto Dominici, Jozef Hladky, Floor Verhoeven, Lukas Radl, Thomas Deixelberger, Stefan Ainetter, Philipp Drescher, Stefan Hauswiesner, Arno Coomans, Giacomo Nazzaro, Konstantinos Vardis, Markus Steinberger

机构 * Huawei Technologies(华为技术有限公司) Graz University of Technology(格拉茨技术大学)

AI总结 DreamAnywhere通过模块化系统实现快速生成和原型设计3D场景,提升新视角合成的连贯性并实现高质量图像输出。

Comments WACV 2026 Oral

URL PDF HTML 收藏
2602.13349 2026-02-17 cs.CV cs.AI

From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models

从提示到生产:利用文本到图像模型自动化安全品牌营销图像

Parmida Atighehchian, Henry Wang, Andrei Kapustin, Boris Lerner, Tiancheng Jiang, Taylor Jensen, Negin Sokhandan

机构 * Amazon Web Services(亚马逊网络服务)

AI总结 本文提出了一种自动化生成品牌安全营销图像的系统,通过文本到图像模型提升图像保真度和人类偏好。

Comments 17 pages, 12 figures, Accepted to IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2602.12983 2026-02-16 cs.CV cs.AI

Detecting Object Tracking Failure via Sequential Hypothesis Testing

通过序列假设检验检测目标跟踪失败

Alejandro Monroy Muñoz, Rajeev Verma, Alexander Timans

机构 * University of Amsterdam(阿姆斯特丹大学)

AI总结 本文提出通过序列假设检验检测目标跟踪失败,提供统计保障的高效方法,适用于实时跟踪系统。

Comments Accepted in WACV workshop "Real World Surveillance: Applications and Challenges, 6th"

URL PDF HTML 收藏
2512.06562 2026-02-13 cs.CV cs.AI

SUGAR: A Sweeter Spot for Generative Unlearning of Many Identities

SUGAR: 为多重身份生成性去学习的更甜位置

Dung Thuy Nguyen, Quang Nguyen, Preston K. Robinette, Eli Jiang, Taylor T. Johnson, Kevin Leach

机构 * Vanderbilt University(范德比大学) Rutgers University(罗格斯大学)

AI总结 SUGAR通过学习个性化替代潜在表示,实现高效移除多个身份的生成性去学习,提升模型保留效用达700%。

Comments IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2602.11024 2026-02-12 cs.CV cs.AI

Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting

密集手术器械计数的链式观察空间推理

Rishikesh Bhyri, Brian R Quaranto, Philip J Seger, Kaity Tung, Brendan Fox, Gene Yang, Steven D. Schwaitzberg, Junsong Yuan, Nan Xi, Peter C W Kim

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校)

AI总结 本文提出Chain-of-Look框架,通过结构化视觉链提升密集手术器械计数的准确性,并引入邻近损失函数和SurgCount-HD数据集,实验证明其在复杂场景中的优越性能。

Comments Accepted to WACV 2026. This version includes additional authors who contributed during the rebuttal phase

URL PDF HTML 收藏
2507.06109 2026-02-12 cs.GR cs.AI cs.CV

LighthouseGS: Indoor Structure-aware 3D Gaussian Splatting for Panorama-Style Mobile Captures

LighthouseGS: 基于室内结构的3D高斯散点式全景式移动捕捉

Seungoh Han, Jaehoon Jang, Hyunsu Kim, Jaeheung Surh, Junhyung Kwak, Hyowon Ha, Kyungdon Joo

机构 * Ulsan National Institute of Science and Technology (UNIST)(乌山国立科学技术研究院)

AI总结 LighthouseGS通过利用室内结构和几何先验,提升单移动设备全景式捕捉的3D点估计和视图合成性能。

Comments WACV 2026

URL PDF HTML 收藏
2408.10831 2026-02-12 cs.CV cs.AI cs.RO

ZebraPose: Zebra Detection and Pose Estimation using only Synthetic Data

ZebraPose: 仅使用合成数据进行斑马检测与姿态估计

Elia Bonetto, Aamir Ahmad

机构 * The International Max Planck Research School for Intelligent Systems(国际马克斯·普朗克智能系统研究学校)

AI总结 ZebraPose通过合成数据集实现斑马检测与姿态估计,解决真实数据获取困难和非分布场景下的模型泛化问题。

Comments 17 pages, 5 tables, 13 figures. Published in WACV 2026

URL PDF HTML 收藏
2602.10160 2026-02-12 cs.CV cs.AI

AD$^2$: Analysis and Detection of Adversarial Threats in Visual Perception for End-to-End Autonomous Driving Systems

AD$^2$:面向端到端自动驾驶系统的视觉感知中对抗威胁的分析与检测

Ishan Sahu, Somnath Hazra, Somak Aditya, Soumyajit Dey

机构 * Indian Institute of Technology Kharagpur(印度理工学院卡里格普尔分校) TCS Research, India(印度塔塔咨询公司研究)

AI总结 AD$^2$通过基于注意力机制的轻量级模型,检测端到端自动驾驶系统中由物理模糊、电磁干扰和数字攻击引起的对抗威胁,提升系统安全性。

Comments Accepted to WACV 2026

URL PDF HTML 收藏
2602.09541 2026-02-11 cs.CV

Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination

Scalpel: 通过混合高斯桥梁实现细粒度注意力激活流形对齐以缓解多模态幻觉

Ziqiang Shi, Rujie Liu, Shanshan Yu, Satoshi Munakata, Koichi Shirahata

机构 * Fujitsu Research & Development Center Co.,LTD.(Fujitsu 研究与开发中心有限公司) Fujitsu Limited(Fujitsu 有限公司)

AI总结 Scalpel通过高斯混合模型和熵最优传输减少多模态幻觉,实现注意力激活流形的细粒度对齐,提升视觉-语言模型的输出一致性。

Comments WACV 2026 (It was accepted in the first round, with an acceptance rate of 6%.)

URL PDF HTML 收藏
2503.20240 2026-02-11 cs.CV

Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models

无条件先验很重要!改进细调扩散模型的条件生成

Prin Phunyaphibarn, Phillip Y. Lee, Jaihoon Kim, Minhyuk Sung

机构 * KAIST(韩国科学技术院)

AI总结 本文提出通过替换无条件噪声提升条件生成质量,验证了使用其他扩散模型进行无条件噪声预测的有效性。

Comments WACV 2026; Project Page: https://unconditional-priors-matter.github.io/

URL PDF HTML 收藏
2602.08996 2026-02-10 cs.CV

Generalizing Sports Feedback Generation by Watching Competitions and Reading Books: A Rock Climbing Case Study

通过观看比赛和阅读书籍来推广体育反馈生成:攀岩案例研究

Arushi Rai, Adriana Kovashka

AI总结 本文通过结合观看比赛和阅读书籍的方法,利用免费网络数据和新指标提升运动反馈生成的泛化能力。

Comments to appear WACV 2026

URL PDF HTML 收藏
2602.08726 2026-02-10 cs.CV

SynSacc: A Blender-to-V2E Pipeline for Synthetic Neuromorphic Eye-Movement Data and Sim-to-Real Spiking Model Training

SynSacc: 一个从Blender到V2E的管道,用于合成神经形态眼动数据和仿真到现实的尖峰模型训练

Khadija Iddrisu, Waseem Shariff, Suzanne Little, Noel OConnor

机构 * Dublin City University(都柏林城市大学) University of Galway(Galway大学)

AI总结 SynSacc利用Blender生成合成数据,通过SNNs训练实现对眼动的高效分类和仿真到现实的尖峰模型训练。

Comments Accepted to the 2nd Workshop on "Event-based Vision in the Era of Generative AI - Transforming Perception and Visual Innovation, IEEE Winter Conference on Applications of Computer Vision (WACV 2026)

URL PDF HTML 收藏
2412.16473 2026-02-10 cs.CV

"ScatSpotter" -- A Dog Poop Detection Dataset

ScatSpotter——一个狗粪便检测数据集

Jon Crall

机构 * Kitware

AI总结 ScatSpotter数据集旨在通过标注狗粪便的多边形图像,研究小型伪装户外废弃物的检测与分割技术,展示了调优DINO模型在该任务中的最佳性能。

Comments Dataset paper, Accepted to the International Workshop on Smart Waste Monitoring (WasteVision) at WACV 2026

URL PDF HTML 收藏
2602.06050 2026-02-09 cs.CL cs.CV

Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering

具有相关性的多上下文对比解码用于检索增强的视觉问答

Jongha Kim, Byungoh Ko, Jeehye Na, Jinsung Yoon, Hyunwoo J. Kim

机构 * Korea University(韩国大学) KAIST(韩国科学技术院) Google Cloud AI(谷歌云人工智能)

AI总结 RMCD通过结合多个相关上下文并抑制无关上下文的负面影响,提升检索增强视觉问答的性能。

Comments WACV 2026

URL PDF HTML 收藏
2509.06165 2026-02-05 cs.CV cs.AI

UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning

UNO:通过对象中心视觉表征学习统一视频场景图生成

Huy Le, Nhat Chung, Tung Kieu, Jingkang Yang, Ngan Le

机构 * FPT Software AI Center(FPT软件AI中心) Aalborg University(奥尔堡大学) Pioneer Centre for AI(先锋人工智能中心) S-Lab, Nanyang Technological University(南洋理工大学S实验室) AICV Lab, University of Arkansas(阿肯色大学AICV实验室)

AI总结 UNO通过统一的对象中心视觉表征学习,实现视频场景图生成的单阶段统一框架,提升效率和泛化能力。

Comments 11 pages, 7 figures. Accepted at WACV 2026

URL PDF HTML 收藏
2601.19136 2026-02-04 cs.CV

TFFM: Topology-Aware Feature Fusion Module via Latent Graph Reasoning for Retinal Vessel Segmentation

TFFM:通过潜在图推理的拓扑感知特征融合模块用于视网膜血管分割

Iftekhar Ahmed, Shakib Absar, Aftar Ahmad Sami, Shadman Sakib, Debojyoti Biswas, Seraj Al Mahmud Mostafa

机构 * Leading University(领先大学) University of Houston(休斯敦大学) University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) Pennsylvania State University(宾夕法尼亚州立大学)

AI总结 TFFM通过拓扑感知特征融合模块提升视网膜血管分割的拓扑连贯性,实现高精度和高可靠性。

Comments Accepted in WACV 2026 @ P2P-workshop as a full paper and selected for oral presentation

URL PDF HTML 收藏
2601.22861 2026-02-03 cs.CV cs.CY cs.ET cs.GR

Under-Canopy Terrain Reconstruction in Dense Forests Using RGB Imaging and Neural 3D Reconstruction

利用RGB成像和神经3D重建进行密林下的地形重建

Refael Sheffer, Chen Pinchover, Haim Zisman, Dror Ozeri, Roee Litman

机构 * Rafael Advanced Defense Systems inc., Israel(拉斐尔先进防御系统公司,以色列) Bar-Ilan University, Israel(巴伊兰大学,以色列)

AI总结 利用RGB图像和神经3D重建技术,实现密林下的地形重建,提供一种低成本、高精度的替代方案。

Comments WACV 2026 CV4EO

URL PDF HTML 收藏
2601.00703 2026-02-03 cs.CV

Efficient Deep Demosaicing with Spatially Downsampled Isotropic Networks

高效深度去马赛克与空间下采样各向同性网络

Cory Fan, Wenchao Zhang

机构 * Cornell University(康奈尔大学) Omnivision(奥米维森) Omnivision Technologies(奥米维森技术)

AI总结 本文提出了一种通过空间下采样提升各向同性网络效率和性能的深度去马赛克方法,并在多种任务中验证了其有效性。

Comments To be published at WVAQ Workshop at WACV. Code @ github.com/cory-fan/jd3net

URL PDF HTML 收藏