arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

至 收录 2109
2601.10054 2026-01-16 cs.CV cs.RO

UEOF: A Benchmark Dataset for Underwater Event-Based Optical Flow

UEOF:用于水下事件驱动光流的基准数据集

Nick Truong, Pritam P. Karmokar, William J. Beksi

机构 * The University of Texas at Arlington(德克萨斯大学阿灵顿分校)

AI总结 本文提出首个水下事件驱动光流基准数据集,通过物理光线追踪生成真实事件数据流,用于评估水下光传输对运动估计的影响。

Comments To be presented at the 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshop on Event-Based Vision in the Era of Generative AI

URL PDF HTML 收藏
2412.16148 2026-01-15 cs.CV

Frequency Is What You Need: Considering Word Frequency When Text Masking Benefits Vision-Language Model Pre-training

频率才是关键:在文本遮蔽时考虑词频有助于视觉-语言模型预训练

Mingliang Liang, Martha Larson

机构 * Institute for Computing and Information Sciences(计算与信息科学研究所) Radboud University(拉德堡德大学)

AI总结 本文提出CLIPF方法,通过考虑词频提升视觉-语言模型预训练效果,实验显示其在减少输入token时表现更优。

Comments Accepted by WACV 2026

URL PDF HTML 收藏
2601.08807 2026-01-14 cs.CV cs.AI

S3-CLIP: Video Super Resolution for Person-ReID

S3-CLIP:基于视频超分辨率的人像重识别

Tamas Endrei, Gyorgy Cserey

机构 * Pázmány Péter Catholic University(帕梅尼·彼得天主教大学)

AI总结 S3-CLIP通过视频超分辨率技术提升轨迹质量,实现人像重识别的性能提升,特别是在跨视图条件下表现优异。

Comments Accepted to the 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW), VReID-XFD Challenge

URL PDF HTML 收藏
2601.08095 2026-01-14 cs.CV

From Prompts to Deployment: Auto-Curated Domain-Specific Dataset Generation via Diffusion Models

从提示到部署:通过扩散模型实现自动定制的领域特定数据集生成

Dongsik Yoon, Jongeun Kim

机构 * HDC LABS(HDC实验室)

AI总结 本文提出通过扩散模型生成领域特定合成数据集的方法,解决预训练模型与现实部署间的分布偏移问题,通过三阶段框架高效构建高质量可部署数据集。

Comments To appear in the Workshop on Synthetic & Adversarial ForEnsics (SAFE), WACV 2026 (oral presentation)

URL PDF HTML 收藏
2504.06856 2026-01-14 cs.CV

CasTex: Cascaded Text-to-Texture Synthesis via Explicit Texture Maps and Physically-Based Shading

CasTex:通过显式纹理映射和基于物理的着色进行级联文本到纹理合成

Mishan Aliev, Dmitry Baranchuk, Kirill Struminsky

机构 * HSE University(俄罗斯莫斯科高等经济大学) Yandex Research(Yandex研究)

AI总结 CasTex通过级联扩散模型实现文本到纹理合成,采用显式纹理参数化提升生成质量,优于现有优化方法。

Comments WACV'2026; v2: camera ready

URL PDF HTML 收藏
2511.17068 2026-01-13 cs.CV cs.AI

ReBrain: Brain MRI Reconstruction from Sparse CT Slice via Retrieval-Augmented Diffusion

ReBrain: 通过检索增强扩散模型从稀疏CT切片重建脑部MRI

Junming Liu, Yifei Sun, Weihua Cheng, Yujin Kang, Yirong Chen, Ding Wang, Guosun Zeng

机构 * Tongji University(同济大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

AI总结 ReBrain通过检索增强扩散模型从稀疏CT切片重建脑部MRI,利用BBDM和ControlNet实现结构连续性,提升稀疏条件下的跨模态重建性能。

Comments 16 pages, 12 figures, 7 tables; Accepted by WACV 2026

URL PDF HTML 收藏
2507.14959 2026-01-13 cs.CV cs.PF

Polymorph: Energy-Efficient Multi-Label Classification for Video Streams on Embedded Devices

Polymorph: 为嵌入式设备上的视频流实现高效多标签分类

Saeid Ghafouri, Mohsen Fayyaz, Xiangchen Li, Deepu John, Bo Ji, Dimitrios Nikolopoulos, Hans Vandierendonck

机构 * Queen’s University Belfast(女王大学贝尔法斯特分校) Microsoft(微软) Virginia Tech(弗吉尼亚理工大学) University College Dublin(都柏林大学)

AI总结 Polymorph通过模块化轻量级LoRA适配器实现嵌入式设备上视频流的高效多标签分类,降低能耗并提升mAP性能。

Comments Accepted at the IEEE/CVF winter conference on applications of computer vision (WACV 2026)

URL PDF HTML 收藏
2503.11742 2026-01-13 cs.CV cs.AI

Safe Vision-Language Models via Unsafe Weights Manipulation

通过不安全权重操作实现安全的视觉-语言模型

Moreno D'Incà, Elia Peruzzo, Xingqian Xu, Humphrey Shi, Nicu Sebe, Massimiliano Mancini

机构 * University of Trento(特伦托大学) NVIDIA(NVIDIA公司) Georgia Tech(佐治亚理工学院)

AI总结 本文提出UWM方法,通过不训练的方式提升视觉-语言模型在不安全查询上的安全性,同时在安全输入上表现更优。

Comments WACV 2026

URL PDF HTML 收藏
2601.06537 2026-01-13 cs.CV

Towards Egocentric 3D Hand Pose Estimation in Unseen Domains

面向未见领域的自体视觉三维手姿态估计

Wiktor Mucha, Michael Wray, Martin Kampel

机构 * Computer Vision Lab, TU Wien(维也纳技术大学计算机视觉实验室) SoftServe Inc.(SoftServe公司) University of Bristol(布里斯托大学)

AI总结 V-HPOT通过虚拟相机空间和自监督优化提升跨领域三维手姿态估计性能,减少71%的平均姿态误差。

Comments Accepted at WACV 2026

URL PDF HTML 收藏
2601.06484 2026-01-13 cs.CV cs.AI

Learning Domain Agnostic Latent Embeddings of 3D Faces for Zero-shot Animal Expression Transfer

学习领域无关的3D人脸潜在嵌入以实现零样本动物表情迁移

Yue Wang, Lawrence Amadi, Xiang Gao, Yazheng Chen, Yuanpeng Liu, Ning Lu, Xianfeng Gu

机构 * Stony Brook University(石英溪大学) Futurewei Technologies(未来讯技术)

AI总结 本文提出了一种零样本框架,通过学习领域无关的3D人脸潜在嵌入,实现人类表情到动物面部的跨物种表情迁移。

Comments WACV 2026 Workshop LENS

URL PDF HTML 收藏
2601.06460 2026-01-13 cs.CV cs.AI cs.CL

Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs

语气至关重要:语言语气对VLMs幻觉影响的研究

Weihao Hong, Zhiyuan Jiang, Bingyu Shen, Xinlei Guan, Yangyi Feng, Meng Xu, Boyang Li

机构 * Department of Computer Science and Technology, Kean University(计算机科学与技术系,凯恩大学) Department of Computer Science and Engineering, University of Notre Dame(计算机科学与工程系,圣母大学)

AI总结 本文研究了提示语气对VLMs幻觉的影响,通过Ghost-100数据集发现幻觉率与提示强度非线性相关,揭示模型在处理结构性强制时的局限性。

Comments 10 pages, 6 figures, WACV Workshop

URL PDF HTML 收藏
2510.25077 2026-01-13 cs.CV eess.IV

Neighborhood Feature Pooling for Remote Sensing Image Classification

邻域特征池化用于遥感图像分类

Fahimeh Orvati Nia, Amirmohammad Mohammadi, Salim Al Kharsa, Pragati Naikare, Zigfried Hampel-Arias, Joshua Peeples

机构 * Dept. of Electrical & Computer Engineering, Texas A&M University, College Station, TX, USA(电子与计算机工程系,德克萨斯A&M大学) Dept. of Computer Science & Engineering, Texas A&M University, College Station, TX, USA(计算机科学与工程系,德克萨斯A&M大学) Los Alamos National Laboratory, Los Alamos, NM, USA(洛斯阿拉莫斯国家实验室)

AI总结 本文提出邻域特征池化方法,通过聚合局部相似性模式提升遥感图像分类性能,实验表明其在多个数据集上均优于传统池化策略。

Comments 10 pages, 4 figures, accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026, 3rd Workshop on Computer Vision for Earth Observation (CV4EO)

URL PDF HTML 收藏
2601.05741 2026-01-12 cs.CV cs.LG

ViTNT-FIQA: Training-Free Face Image Quality Assessment with Vision Transformers

ViTNT-FIQA: 基于视觉变换器的无训练面部图像质量评估

Guray Ozgur, Eduarda Caldeira, Tahar Chettaoui, Jan Niklas Kolf, Marco Huber, Naser Damer, Fadi Boutros

AI总结 ViTNT-FIQA通过测量视觉变换器中间块补丁嵌入演变的稳定性,实现无训练的面部图像质量评估,具有高效计算和广泛适用性。

Comments Accepted at WACV Workshops

URL PDF HTML 收藏
2601.04184 2026-01-12 cs.MM

Transforming Video Subjective Testing with Training, Engagement, and Real-Time Feedback

通过训练、参与和实时反馈转变视频主观测试

Kumar Rahul, Sriram Sethuraman, Andrew Segall, Yixu Chen

AI总结 本文提出一种通过训练、参与和实时反馈改进视频主观测试的新方法,提升数据质量和评分单调性。

Comments Accepted at 5th Workshop on Image/Video/Audio Quality Assessment in Computer Vision, VLM and Diffusion Model (WVAQ), at IEEE/CVF WACV 2026

URL PDF HTML 收藏
2601.04798 2026-01-09 cs.CV

Detector-Augmented SAMURAI for Long-Duration Drone Tracking

增强检测器的SAMURAI用于长时间无人机跟踪

Tamara R. Lenhard, Andreas Weinmann, Hichem Snoussi, Tobias Koch

机构 * Institute for the Protection of Terrestrial Infrastructures, German Aerospace Center (DLR)(地面基础设施保护研究所,德国航空航天中心(DLR)) ACIDA Lab, Technical University of Applied Sciences Würzburg-Schweinfurt(应用技术大学施魏尔堡-施维恩富特学院ACIDA实验室) Data Science Institute, European University of Technology(欧洲技术大学数据科学研究所) LIST3N, Université de Technologie de Troyes(图卢兹理工大学LIST3N)

AI总结 本文提出增强检测器的SAMURAI模型,用于提升无人机在复杂城市环境中的长时间跟踪鲁棒性,显著提高了成功率并降低了误检率。

Comments Accepted at the WACV 2026 Workshop on "Real World Surveillance: Applications and Challenges"

URL PDF HTML 收藏
2512.17226 2026-01-09 cs.CV

Robust Scene Coordinate Regression via Geometrically-Consistent Global Descriptors

通过几何一致的全局描述符实现鲁棒的场景坐标回归

Son Tung Nguyen, Alejandro Fontan, Michael Milford, Tobias Fischer

机构 * Queensland University of Technology(昆士兰理工大学)

AI总结 本文提出一种通过几何一致的全局描述符提升场景坐标回归鲁棒性的方法,无需人工标注即可在多样环境中实现高效定位。

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2411.05633 2026-01-09 cs.CV cs.AI cs.RO

SynDroneVision: A Synthetic Dataset for Image-Based Drone Detection

SynDroneVision:用于图像-based无人机检测的合成数据集

Tamara R. Lenhard, Andreas Weinmann, Kai Franke, Tobias Koch

机构 * Institute for the Protection of Terrestrial Infrastructures, German Aerospace Center (DLR)(地面基础设施保护研究所,德国航空航天中心(DLR)) Working Group Algorithms for Computer Vision, Imaging and Data Analysis, University of Applied Sciences Darmstadt(计算机视觉、成像与数据分析算法工作组,达姆施塔特应用科学大学)

AI总结 SynDroneVision通过合成数据提升无人机检测性能,降低真实数据采集成本

Journal ref 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

URL PDF HTML 收藏
2601.02908 2026-01-07 cs.CV cs.AI cs.LG

TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors

TA-Prompting: 通过时间锚点增强视频大语言模型以实现密集视频描述

Wei-Yuan Cheng, Kai-Po Chang, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国家交通大学通信工程研究所) NVIDIA(NVIDIA公司)

AI总结 TA-Prompting通过引入时间锚点提升视频大语言模型的密集视频描述能力,有效解决时间边界识别问题,提升时间感知视频事件理解性能。

Comments 8 pages for main paper (exclude citation pages), 6 pages for appendix, totally 10 figures 7 tables and 2 algorithms. The paper is accepted by WACV 2026

Journal ref IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2411.11917 2026-01-07 cs.CV

FCC: Fully Connected Correlation for One-Shot Segmentation

FCC:全连接相关性用于一次学习分割

Seonghyeon Moon, Haein Kong, Muhammad Haris Khan, Mubbasir Kapadia, Yuewei Lin

机构 * Roblox Brookhaven National Laboratory(布鲁克海文国家实验室) Rutgers University(罗格斯大学) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 FCC通过整合支持和查询特征间的像素级相关性,提升少样本分割的性能,有效解决支持掩码的局限性。

Comments WACV 2026

URL PDF HTML 收藏
2601.02315 2026-01-06 cs.CV

Prithvi-Complimentary Adaptive Fusion Encoder (CAFE): unlocking full-potential for flood inundation mapping

普里提维-互补自适应融合编码器(CAFE):解锁洪水淹没制图的全部潜力

Saurabh Kaushik, Lalit Maurya, Beth Tellman

机构 * Center for Sustainability and the Global Environment (SAGE), University of Wisconsin–Madison(可持续性与全球环境中心(SAGE),威斯康星大学麦迪逊分校) Portsmouth AI and Data Science Centre (PAIDS), School of Computing, University of Portsmouth(波特兰人工智能与数据科学中心(PAIDS),计算学院,波特兰大学)

AI总结 普里提维-互补自适应融合编码器(CAFE)通过融合多通道多模态数据提升洪水制图的分割性能。

Comments Accepted at CV4EO Workshop @ WACV 2026

URL PDF HTML 收藏
2601.02289 2026-01-06 cs.CV

Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery

基于排名的地理正则化:重新审视多光谱遥感图像的对比自监督学习

Tom Burgert, Leonard Hackel, Paolo Rota, Begüm Demir

机构 * BIFOLD TU Berlin(柏林技术大学) University of Trento(特伦托大学)

AI总结 本文提出GeoRank,一种改进对比自监督学习的地理正则化方法,通过优化球面距离嵌入地理关系,提升多光谱遥感图像的特征学习效果。

Comments accepted for publication at IEEE/CVF Winter Conference on Applications of Computer Vision

URL PDF HTML 收藏
2601.01963 2026-01-06 cs.CV cs.LG

Forget Less by Learning Together through Concept Consolidation

通过概念整合实现学习共存以减少遗忘

Arjun Ramesh Kaushik, Naresh Kumar Devulapally, Vishnu Suresh Lokhande, Nalini Ratha, Venu Govindaraju

机构 * University at Buffalo, SUNY(布法罗大学)

AI总结 本文提出FL2T框架,通过跨概念学习模块减少灾难性遗忘,提升概念保留与迁移效果。

Comments Accepted at WACV-26

URL PDF HTML 收藏
2601.01914 2026-01-06 cs.CV

Learning Action Hierarchies via Hybrid Geometric Diffusion

通过混合几何扩散学习动作层次结构

Arjun Ramesh Kaushik, Nalini K. Ratha, Venu Govindaraju

AI总结 本文提出HybridTAS框架,通过混合欧几里得和双曲几何提升时间动作分割的层次结构建模能力,实验表明其在多个数据集上达到最优性能。

Comments Accepted at WACV-26

URL PDF HTML 收藏
2405.20610 2026-01-06 cs.CV

PrevMatch: Revisiting and Maximizing Temporal Knowledge in Semi-Supervised Semantic Segmentation

PrevMatch: 重新审视并最大化半监督语义分割中的时间知识

Wooseok Shin, Hyun Joon Park, Jin Sob Kim, Juan Yun, Se Hong Park, Sung Won Han

机构 * Department of Industrial and Management Engineering, Korea University(韩国大学工业与管理工程系)

AI总结 PrevMatch通过最大化训练过程中的时间知识,提升半监督语义分割的性能和可扩展性。

Comments To appear in WACV 2026. Code: https://github.com/wooseok-shin/PrevMatch

URL PDF HTML 收藏
2405.18770 2026-01-06 cs.CV cs.AI cs.IR

Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships

通过利用一对一关系进行多模态对抗防御以提升视觉语言模型

Futa Waseda, Antonio Tejero-de-Pablos, Isao Echizen

机构 * The University of Tokyo(东京大学) CyberAgent National Institute of Informatics(信息处理研究所)

AI总结 本文提出多模态对抗训练方法,通过利用图像与文本之间的一对多关系提升视觉语言模型的对抗鲁棒性。

Comments WACV 2026 Accepted. Code available at https://github.com/CyberAgentAILab/multimodal-adversarial-training

URL PDF HTML 收藏
2601.01103 2026-01-06 cs.CV eess.IV

Histogram Assisted Quality Aware Generative Model for Resolution Invariant NIR Image Colorization

基于直方图的质量感知生成模型用于分辨率不变的近红外图像着色

Abhinav Attri, Rajeev Ranjan Dwivedi, Samiran Das, Vinod Kumar Kurmi

机构 * Indian Institute of Science Education and Research Bhopal(印度科学教育与研究学院博帕尔)

AI总结 HAQAGen通过结合直方图匹配、SPADE和Mamba网络,实现分辨率不变的近红外图像到RGB的高质量着色,提升纹理和颜色的真实感。

Comments Accepted at WACV 2026

URL PDF HTML 收藏
2601.01062 2026-01-06 cs.LG cs.AI cs.CV

SPoRC-VIST: A Benchmark for Evaluating Generative Natural Narrative in Vision-Language Models

SPoRC-VIST:评估视觉语言模型生成自然叙述能力的基准

Yunlin Zeng

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 SPoRC-VIST基准通过合成到现实训练策略评估视觉语言模型生成多说话者播客对话的自然性和深度。

Comments 14 pages, 3 figures. Accepted to WVAQ 2026, WACV 2026

URL PDF HTML 收藏
2601.01024 2026-01-06 cs.CV cs.AI cs.IR

ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval

ITSELF: 基于注意力引导的细粒度对齐用于视觉-语言检索

Tien-Huy Nguyen, Huu-Loc Tran, Thanh Duc Ngo

机构 * University of Information Technology(信息技术大学) Vietnam National University(越南国家大学)

AI总结 ITSELF通过基于注意力引导的细粒度对齐框架,在视觉-语言检索任务中实现最先进的性能和跨数据集泛化能力。

Comments Accepted at WACV Main Track 2026

URL PDF HTML 收藏
2601.00943 2026-01-06 cs.CV

PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education

PhyEduVideo: 一个用于评估物理教育中文本到视频模型的基准

Megha Mariam K. M, Aditya Arun, Zakaria Laskar, C. V. Jawahar

机构 * IIIT Hyderabad(印度海得拉巴理工学院) Adobe MDSR IISER Thiruvananthapuram(泰米尔纳德邦土著科学研究所)

AI总结 PhyEduVideo基准评估文本到视频模型在物理教育中的表现,揭示其在概念准确性与视觉质量间的差距,推动更精准的教育视频生成。

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

URL PDF HTML 收藏
2601.00222 2026-01-05 cs.CV

LooC: Effective Low-Dimensional Codebook for Compositional Vector Quantization

LooC: 有效的低维代码书用于组合向量量化

Jie Li, Kwan-Yee K. Wong, Kai Han

机构 * The University of Hong Kong(香港大学)

AI总结 LooC通过低维代码书提升组合向量量化的效率与性能,实现更紧凑的代码书和更优的特征近似。

Comments The IEEE/CVF Winter Conference on Applications of Computer Vision 2026

URL PDF HTML 收藏