arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

2026-04-24 至 2026-04-24 共收录 9
2604.21453 2026-04-24 cs.CV

Instance-level Visual Active Tracking with Occlusion-Aware Planning

实例级视觉主动跟踪与遮挡感知规划

Haowei Sun, Kai Zhou, Hao Gao, Shiteng Zhang, Jinwu Hu, Xutao Wen, Qixiang Ye, Mingkui Tan

机构 * South China University of Technology(南方科技大学) Pazhou Laboratory(Pazhou实验室) Key Laboratory of Big Data and Intelligent Robot, Ministry of Education(教育部大数据与智能机器人重点实验室) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 本文提出OA-VAT方法,通过实例感知原型初始化、在线原型增强跟踪和遮挡感知轨迹规划模块,解决遮挡和相似干扰问题,实现高精度实时跟踪。

Comments CVPR 2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21127 2026-04-24 cs.CV

HyperFM: An Efficient Hyperspectral Foundation Model with Spectral Grouping

HyperFM:一种高效的超光谱基础模型与光谱分组

Zahid Hassan Tushar, Sanjay Purushotham

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩分校)

AI总结 本文提出HyperFM,一种高效超光谱基础模型,通过光谱分组和混合参数分解,提升对光谱空间关系的建模能力,同时降低计算成本,在四个大气云特性检索任务中表现优异,并发布HyperFM250K数据集支持进一步研究。

Comments 15 pages, 8 figures, to be published in CVPR 2026 findings, Code and data are publicly available on https://github.com/umbc-sanjaylab/HyperFM

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21119 2026-04-24 cs.CV cs.AI cs.SD

Materialistic RIR: Material Conditioned Realistic RIR Generation

物质导向的RIR:基于材料的现实RIR生成

Mahnoor Fatima Saad, Sagnik Majumder, Kristen Grauman, Ziad Al-Halah

机构 * University of Utah(犹他大学) UT Austin(得克萨斯大学奥斯汀分校)

AI总结 本文提出一种基于材料的RIR生成方法,通过分离空间和材料影响,提升生成声学的真实感和材料敏感性,实验结果显示在声学和材料指标上均有显著提升。

Comments Accepted to CVPR 2026 Findings. Project page: https://mahnoor-fatima-saad.github.io/MatRIR.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21104 2026-04-24 cs.CV cs.LG

Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance

预训练在哪里?探究预训练数据多样性如何影响地理空间基础模型性能

Amandeep Kaur, Mirali Purohit, Gedeon Muhawenayo, Esther Rolf, Hannah Kerner

机构 * Arizona State University(亚利桑那州立大学) University of Colorado Boulder(科罗拉多大学博尔德分校)

AI总结 本文研究预训练数据地理组成对模型下游性能的影响,发现欧洲预训练数据在全局和本地评估中表现最佳,发现光谱多样性与性能强相关。

Comments Accepted at EarthVision workshop, CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20878 2026-04-24 cs.CL cs.CV cs.LG eess.IV

AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models

AITP:通过多模态大语言模型进行交通事故责任分配

Zijin Zhou, Songan Zhang

机构 * Global Institute of Future Technology(未来技术全球研究院)

AI总结 本文提出AITP模型,结合多模态链式推理和法律知识检索,解决交通事故责任分配问题,并构建DecaTARA基准测试集,验证模型在责任分配、事故检测和理解任务中的领先性能。

Journal ref CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10275 2026-04-24 cs.CV

FastSHADE: Fast Self-augmented Hierarchical Asymmetric Denoising for Efficient inference on mobile devices

FastSHADE:快速自增强分层非对称去噪用于移动设备上的高效推理

Nikolay Falaleev

机构 * Fanis London, UK(Fanis伦敦大学)

AI总结 本文提出FastSHADE,一种轻量级U-Net风格网络,用于移动设备上的实时高质量去噪。通过非对称频率去噪块和空间门控上采样器提升效率与图像质量,实现高效速度-保真度平衡。

Comments To appear in the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12845 2026-04-24 cs.CV

Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

多模态蛋白质语言模型用于酶动力学参数:从底物识别到构象适应

Fei Wang, Xinye Zheng, Kun Li, Yanyan Wei, Yuxin Liu, Ganpeng Hu, Tong Bao, Jingwen Yang

机构 * School of Computer Science and Information Engineering(计算机科学与信息工程学院) Institute of Artificial Intelligence(人工智能研究院) CVLab, College of Information Technology(CV实验室,信息学院) Intelligent Interconnected Systems Laboratory of Anhui Province(安徽省智能互联系统实验室) School of Food Biological Engineering(食品生物工程学院)

AI总结 本文提出多阶段多模态条件建模方法,通过ERBA模块在蛋白质语言模型中注入跨模态信息,提升酶动力学参数预测的准确性与生物合理性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02409 2026-04-24 cs.CV

Catalyst: Out-of-Distribution Detection via Elastic Scaling

催化剂:通过弹性缩放进行分布外检测

Abid Hassan, Tuan Ngo, Saad Shafiq, Nenad Medvidovic

机构 * University Southern California(南加州大学)

AI总结 本文提出Catalyst框架,通过利用预池化特征图的原始通道统计信息,改进分布外检测的性能,显著降低误报率。

Comments Accepted at Conference on Computer Vision and Pattern Recognition (CVPR) 2026. arXiv admin note: text overlap with arXiv:2601.22703

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18457 2026-04-24 cs.CV cs.LG

VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models

VFM-VAE:视觉基础模型可以作为潜在扩散模型的良好分词器

Tianci Bi, Xiaoyi Zhang, Yan Lu, Nanning Zheng

机构 * State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学) Microsoft Research Asia(微软亚洲研究院)

AI总结 本文提出VFM-VAE,利用冻结的视觉基础模型作为潜在扩散模型的分词器,通过设计新解码器提升图像重建能力,实现分词器与扩散模型的协同优化,提升训练效率与性能。

Comments Accepted at CVPR 2026. Code and models available at: https://github.com/tianciB/VFM-VAE

详情

展开后加载摘要…

URL PDF HTML 收藏