arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

2026-01-06 至 2026-01-06 共收录 10
2601.02315 2026-01-06 cs.CV

Prithvi-Complimentary Adaptive Fusion Encoder (CAFE): unlocking full-potential for flood inundation mapping

普里提维-互补自适应融合编码器(CAFE):解锁洪水淹没制图的全部潜力

Saurabh Kaushik, Lalit Maurya, Beth Tellman

机构 * Center for Sustainability and the Global Environment (SAGE), University of Wisconsin–Madison(可持续性与全球环境中心(SAGE),威斯康星大学麦迪逊分校) Portsmouth AI and Data Science Centre (PAIDS), School of Computing, University of Portsmouth(波特兰人工智能与数据科学中心(PAIDS),计算学院,波特兰大学)

AI总结 普里提维-互补自适应融合编码器(CAFE)通过融合多通道多模态数据提升洪水制图的分割性能。

Comments Accepted at CV4EO Workshop @ WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02289 2026-01-06 cs.CV

Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery

基于排名的地理正则化:重新审视多光谱遥感图像的对比自监督学习

Tom Burgert, Leonard Hackel, Paolo Rota, Begüm Demir

机构 * BIFOLD TU Berlin(柏林技术大学) University of Trento(特伦托大学)

AI总结 本文提出GeoRank,一种改进对比自监督学习的地理正则化方法,通过优化球面距离嵌入地理关系,提升多光谱遥感图像的特征学习效果。

Comments accepted for publication at IEEE/CVF Winter Conference on Applications of Computer Vision

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01963 2026-01-06 cs.CV cs.LG

Forget Less by Learning Together through Concept Consolidation

通过概念整合实现学习共存以减少遗忘

Arjun Ramesh Kaushik, Naresh Kumar Devulapally, Vishnu Suresh Lokhande, Nalini Ratha, Venu Govindaraju

机构 * University at Buffalo, SUNY(布法罗大学)

AI总结 本文提出FL2T框架,通过跨概念学习模块减少灾难性遗忘,提升概念保留与迁移效果。

Comments Accepted at WACV-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01914 2026-01-06 cs.CV

Learning Action Hierarchies via Hybrid Geometric Diffusion

通过混合几何扩散学习动作层次结构

Arjun Ramesh Kaushik, Nalini K. Ratha, Venu Govindaraju

AI总结 本文提出HybridTAS框架,通过混合欧几里得和双曲几何提升时间动作分割的层次结构建模能力,实验表明其在多个数据集上达到最优性能。

Comments Accepted at WACV-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20610 2026-01-06 cs.CV

PrevMatch: Revisiting and Maximizing Temporal Knowledge in Semi-Supervised Semantic Segmentation

PrevMatch: 重新审视并最大化半监督语义分割中的时间知识

Wooseok Shin, Hyun Joon Park, Jin Sob Kim, Juan Yun, Se Hong Park, Sung Won Han

机构 * Department of Industrial and Management Engineering, Korea University(韩国大学工业与管理工程系)

AI总结 PrevMatch通过最大化训练过程中的时间知识,提升半监督语义分割的性能和可扩展性。

Comments To appear in WACV 2026. Code: https://github.com/wooseok-shin/PrevMatch

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18770 2026-01-06 cs.CV cs.AI cs.IR

Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships

通过利用一对一关系进行多模态对抗防御以提升视觉语言模型

Futa Waseda, Antonio Tejero-de-Pablos, Isao Echizen

机构 * The University of Tokyo(东京大学) CyberAgent National Institute of Informatics(信息处理研究所)

AI总结 本文提出多模态对抗训练方法,通过利用图像与文本之间的一对多关系提升视觉语言模型的对抗鲁棒性。

Comments WACV 2026 Accepted. Code available at https://github.com/CyberAgentAILab/multimodal-adversarial-training

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01103 2026-01-06 cs.CV eess.IV

Histogram Assisted Quality Aware Generative Model for Resolution Invariant NIR Image Colorization

基于直方图的质量感知生成模型用于分辨率不变的近红外图像着色

Abhinav Attri, Rajeev Ranjan Dwivedi, Samiran Das, Vinod Kumar Kurmi

机构 * Indian Institute of Science Education and Research Bhopal(印度科学教育与研究学院博帕尔)

AI总结 HAQAGen通过结合直方图匹配、SPADE和Mamba网络,实现分辨率不变的近红外图像到RGB的高质量着色,提升纹理和颜色的真实感。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01062 2026-01-06 cs.LG cs.AI cs.CV

SPoRC-VIST: A Benchmark for Evaluating Generative Natural Narrative in Vision-Language Models

SPoRC-VIST:评估视觉语言模型生成自然叙述能力的基准

Yunlin Zeng

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 SPoRC-VIST基准通过合成到现实训练策略评估视觉语言模型生成多说话者播客对话的自然性和深度。

Comments 14 pages, 3 figures. Accepted to WVAQ 2026, WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01024 2026-01-06 cs.CV cs.AI cs.IR

ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval

ITSELF: 基于注意力引导的细粒度对齐用于视觉-语言检索

Tien-Huy Nguyen, Huu-Loc Tran, Thanh Duc Ngo

机构 * University of Information Technology(信息技术大学) Vietnam National University(越南国家大学)

AI总结 ITSELF通过基于注意力引导的细粒度对齐框架,在视觉-语言检索任务中实现最先进的性能和跨数据集泛化能力。

Comments Accepted at WACV Main Track 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00943 2026-01-06 cs.CV

PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education

PhyEduVideo: 一个用于评估物理教育中文本到视频模型的基准

Megha Mariam K. M, Aditya Arun, Zakaria Laskar, C. V. Jawahar

机构 * IIIT Hyderabad(印度海得拉巴理工学院) Adobe MDSR IISER Thiruvananthapuram(泰米尔纳德邦土著科学研究所)

AI总结 PhyEduVideo基准评估文本到视频模型在物理教育中的表现,揭示其在概念准确性与视觉质量间的差距,推动更精准的教育视频生成。

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏