arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

共收录 2122 篇
2609.37314 2026-09-30 cs.CV 新提交

FLASH: A "Generate Once, Synthesize Many" Framework for Synthetic Anomaly Generation in Industrial Anomaly Detection

FLASH:工业异常检测中合成异常生成的“一次生成,多次合成”框架

Abhay Kumar Das, Rajesh Gangireddy, Ashwin Vaidya, Samet Akcay

机构 * Silicon University(硅谷大学) ; Intel(英特尔)

AI总结 FLASH提出“一次生成,多次合成”框架,解耦缺陷生成与异常合成,利用VLM和图像生成模型提取可复用缺陷补丁,结合OBS和MRSP高效合成多样异常,在MVTec AD 2上以78.1% F1接近真实异常性能,速度提升11.95倍。

Comments Submitted to WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.29456 2026-09-25 cs.CV 新提交

Dense Coverage, Sparse Refinement: Byte-Constrained Cooperative Perception

密集覆盖,稀疏细化:字节受限的协同感知

Melih Yazgan, Timon Müller, J. Marius Zöllner

机构 * FZI Research Center for Information Technology(FZI信息技术研究中心) ; Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

AI总结 针对V2X带宽限制下的协同感知,提出覆盖-细化方法:全图粗层加高价值补丁细化,任务感知选择器分配预算,在DAIR-V2X和OPV2V上以千字节级负载实现高精度。

Comments Accepted at WACV 2027 (first-round acceptance)

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.23427 2026-09-24 cs.CV 版本更新

RSPDBench: Benchmarking Vision Foundation Models on Earth Observation Tasks Under Physically Grounded Remote-Sensing Product Degradations

RSPDBench:在物理接地遥感产品退化条件下对地球观测任务视觉基础模型的基准测试

Tanjim Bin Faruk, Khondaker Masfiq Reza, Shrideep Pallickara, Sangmi Lee Pallickara

机构 * Colorado State University(科罗拉多州立大学)

AI总结 该研究提出物理接地的遥感产品退化基准RSPDBench,评估视觉基础模型在真实EO产品缺陷下的鲁棒性,发现退化敏感性具有结构性,复合退化可导致高达38个百分点的额外性能下降。

Comments Accepted to WACV 2027 (Round 1)

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.23003 2026-09-22 cs.CV cs.RO 新提交

M3GA-Wild: A Large-Scale Dataset and Benchmark for Multi-Modal Multi-session Ground-to-Aerial Place Recognition in Forests

M3GA-Wild:用于森林中多模态、多会话地空地点识别的大规模数据集与基准

Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Milad Ramezani

机构 * Queensland University of Technology (QUT)(昆士兰科技大学) ; CSIRO Robotics(澳大利亚联邦科学与工业研究组织机器人部门)

AI总结 M3GA-Wild是首个森林多模态多会话地空地点识别基准,含36公里地面和370公顷航空数据,揭示LiDAR优于视觉及跨模态对齐挑战。

Comments Accepted for publication in WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.20962 2026-09-21 cs.CV cs.MM 新提交

MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction

MemeTAG:通过标签嵌入重建实现关键词驱动的模因分类

Akshit Sharma, Prashant W. Patil

机构 * Indian Institute of Technology Guwahati(印度理工学院古瓦哈提分校)

AI总结 MemeTAG提出双目标框架,利用视觉-语言模型生成关键词并经ATIN模块聚合为语义嵌入,通过辅助重建损失对齐视觉与文本特征,在多个数据集上超越现有方法。

Comments 10 pages, 3 figures; published in the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026, pp. 7679-7688

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14962 2026-09-18 cs.LG cs.AI cs.CV 版本更新

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation

用于文本到图像生成中代表性多样性的多轴最大@K强化学习

Ku Onoda, Paavo Parmas, Hiroki Furuta, Soichiro Nishimori, Yuta Oshima, Shohei Taniguchi, Yutaka Matsuo

机构 * The University of Tokyo(东京大学)

AI总结 研究文本到图像生成中样本模式覆盖不足问题,提出多轴最大@K强化学习目标,通过特定信用分配机制改善覆盖,在感知外观公平性评估中提高公平分数,且保持图像质量和文本对齐。

Comments Accepted at WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.16946 2026-09-17 cs.CV 版本更新

High-Fidelity Video Quality Assessment with VQA-Specific Saliency

高保真视频质量评估与VQA特定显著性

Hakan Emre Gedik, Shashank Gupta, Alan Bovik

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) ; University of Colorado Boulder(科罗拉多大学博尔德分校)

AI总结 提出HFVQA框架,利用固定大小时空补丁与VQA特定显著性,在保留高保真线索的同时仅处理12%补丁,实现无参考视频质量评估的最先进性能。

Comments Accepted to WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.17458 2026-09-16 cs.CV cs.LG 新提交

Tables Decoded: DELTA for Structure, TARQA for Understanding

表格解码:DELTA 用于结构,TARQA 用于理解

Jahanvi Rajput, Dhruv Kudale, Saikiran Kasturi, Utkarsh Verma, Ganesh Ramakrishnan

机构 * Indian Institute of Technology Bombay(印度理工学院孟买分校) ; BharatGen

AI总结 针对表格理解,提出基于结构化文本的DELTA和TARQA,分别处理结构识别和问答,在多个基准上达到先进性能,并验证了多语言鲁棒性。

Comments Accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.14466 2026-09-15 cs.CV cs.RO 新提交

DynEoMT: Learning Object Dynamicity from Online Segmentation Queries

DynEoMT:从在线分割查询中学习物体动态性

Calvin Galagain, Martyna Poreba, François Goulette

机构 * Université Paris-Saclay(巴黎-萨克雷大学) ; CEA(法国原子能委员会) ; List(List研究所) ; ENSTA Paris(巴黎高科先进技术学校) ; Institut Polytechnique de Paris(巴黎综合理工学院)

AI总结 DynEoMT是一个在线框架,通过传播查询为视频分割区域预测动态性,无需光流或深度,在多个基准上取得高平衡准确率,同时保持分割性能。

Comments Submitted to the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.13258 2026-09-15 cs.CV cs.CL cs.IR 新提交

Interpretable Temporal Video Reasoning with EventGraph and EventField

可解释的时间视频推理:EventGraph 与 EventField

Durgendra Narayan Singh

AI总结 提出基于EventGraph和EventField的结构化时间视频推理流水线,在EPIC-KITCHENS子集上准确率达0.98,优于字幕基线和VLM,兼顾性能与可解释性。

Comments Submitted to WACV 2027. Preprint; 9 pages, 7 figures. Copyright may be transferred without notice

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07776 2026-09-15 cs.CV 版本更新

GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring

GorillaWatch: 一种用于野外大猩猩重识别与种群监测的自动化系统

Maximilian Schall, Felix Leonard Knöfel, Noah Elias König, Jan Jonas Kubeler, Maximilian von Klinski, Joan Wilhelm Linnemann, Xiaoshi Liu, Iven Jelle Schlegelmilch, Ole Woyciniuk, Alexandra Schild, Dante Wasmuht, Magdalena Bermejo Espinet, German Illera Basas, Gerard de Melo

机构 * Hasso Plattner Institute(哈索普拉特纳研究所) ; Conservation X Labs(保护X实验室) ; Sabine Plattner African Charities(萨宾·普拉特纳非洲慈善机构)

AI总结 GorillaWatch通过引入三个新数据集和端到端流程,实现了野外大猩猩的自动化重识别与种群监测,结合自监督预训练和可微分适应技术提升模型有效性。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06433 2026-09-14 cs.CV

Diagnose Like A REAL Pathologist: An Uncertainty-Focused Approach for Trustworthy Multi-Resolution Multiple Instance Learning

像真正的病理科医生一样诊断:一种以不确定性为核心的多分辨率多实例学习方法

Sungrae Hong, Sol Lee, Jisu Shin, Jiwon Jeong, Mun Yong Yi

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

AI总结 本文提出UFC-MIL方法,通过多分辨率图像和不确定性校准,提升多实例学习的诊断可靠性。

Comments Accepted by IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

Journal ref 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.08660 2026-09-09 cs.CV 新提交

CoordFormer: Give Me Any Coordinates and I Will Give You Labels

CoordFormer:给我任意坐标,我便给你标签

Iacopo Curti, Pierluigi Zama Ramirez, Alioscia Petrelli, Luigi Di Stefano

机构 * CVLab, University of Bologna(博洛尼亚大学CV实验室) ; Ca’ Foscari University of Venice(威尼斯大学) ; SINA, company(SINA公司)

AI总结 CoordFormer提出基于坐标的解码器与局部交叉注意力机制,结合边缘聚焦策略,实现低内存、高精度的超高分辨率图像语义分割,在MaSS13K等基准上达到最优性能。

Comments Accepted at WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.01041 2026-09-02 cs.CV cs.AI cs.LG 新提交

ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives

ViTAMINS:使用合成难负样本训练自监督视觉Transformer的实证研究

Nikos Giakoumoglou, Andreas Floros, Kleanthis-Marios Papadopoulos, Tania Stathaki

机构 * Imperial College London(帝国理工学院)

AI总结 本研究提出ViTAMINS方法,将合成难负样本用于视觉Transformer预训练,经多任务基准测试,该方法性能优于竞争模型且资源效率更高,推动对比学习成为主流生成式与自蒸馏方法的替代方案。

Comments WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07632 2026-09-02 q-bio.QM cs.CV eess.IV 版本更新

JUMP-lite: Compact, reproducible benchmarking of cell representations

JUMP-lite:紧凑、可复现的细胞表征基准测试

Alán F. Muñoz, Johan Fredin Haslum, Runxi Shen, Anne E. Carpenter, Shantanu Singh

机构 * Broad Institute of MIT and Harvard(麻省理工学院与哈佛大学博德研究所)

AI总结 研究针对JUMP数据集体积过大导致细胞表征基准测试难以开展的问题,提出JUMP-lite压缩子集与Nahual框架,测试5种表征方法并验证压缩保留下游信号,为相关基准测试提供基础。

Comments Submitted to WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11769 2026-09-02 cs.CV

From Pixels to Purchase: Building and Evaluating a Taxonomy-Decoupled Visual Search Engine for Home Goods E-commerce

从像素到购买:构建和评估一个去分类的视觉搜索引擎用于家居商品电商

Cheng Lyu, Jingyue Zhang, Ryan Maunu, Mengwei Li, Vinny DeGenova, Yuanli Pei

机构 * Wayfair

AI总结 本文提出了一种去分类的视觉搜索引擎,通过无分类区域提案和统一嵌入提升搜索灵活性,并引入LLM-as-a-Judge框架实现零样本评估,从而提高家居电商的检索质量和客户参与度。

Journal ref 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW), pp. 1582-1590, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22059 2026-08-25 eess.IV cs.AI cs.CV 新提交

CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders

CRS-Bench:面向医学图像编码器的参考相对可靠性基准

Xingtao Lin, Hangqi Ren, Caiwan Sun, You Chen

机构 * Vanderbilt University Medical Center(范德堡大学医学中心) ; Vanderbilt University(范德堡大学)

AI总结 本研究提出CRS-Bench基准,通过评估15个预训练医学图像编码器的多轴可靠性,结合临床可靠性评分,发现部分编码器排序与AUROC结果反转,确定PanDerm等为稳定领先层级,为医学编码器选择提供更全面框架。

Comments 10 pages, 7 figures. Submitted to WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02846 2026-08-21 cs.CV

Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?

瞬间动作预见:多模态线索能替代视频到何种程度?

Manuel Benavent-Lledo, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez

机构 * Universidad de Alicante(阿利坎特大学) ; Foundation for Research and Technology-Hellas(希腊基础研究与技术基金会) ; University of Crete(克里特大学) ; Hellenic Mediterranean University(希腊地中海大学)

AI总结 AAG通过结合单帧RGB特征与深度线索及先前动作信息,实现了多模态单帧动作预见,能与视频聚合基线和先进方法在教学活动数据集上竞争。

Comments Accepted in WACV 2026 - Applications Track

Journal ref 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01781 2026-08-20 cs.CV cs.AI cs.LG

Subimage Overlap Prediction: Task-Aligned Self-Supervised Pretraining For Semantic Segmentation In Remote Sensing Imagery

子图像重叠预测:面向遥感图像语义分割的任务对齐自监督预训练

Lakshay Sharma, Alex Marin

机构 * Instacart ; New York University(纽约大学) ; University of Washington(华盛顿大学)

AI总结 本文提出子图像重叠预测任务,通过少量预训练数据提升遥感图像语义分割性能,实现更快收敛和同等或更优的mIoU表现。

Comments Accepted at CV4EO Workshop at WACV 2026

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops, 2026, pp. 1414-1423

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15349 2026-08-18 cs.CV cs.AI eess.IV 新提交

ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution

ENAF:用于大图像超分辨率的自适应补丁融合多出口网络

Duong M. Nguyen, Tuan Nghia Nguyen, Xuan Truong Nguyen

机构 * Hanoi University of Science and Technology(河内理工大学) ; Seoul National University(首尔大学)

AI总结 ENAF是带自适应补丁融合的SISR动态网络,通过嵌入小型PSNR估计网络优化补丁分配,在常用数据集上结合主流骨干网络可提升SISR的质量-复杂度权衡效果

Comments Accepted at WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04798 2026-08-18 cs.CV 版本更新

Detector-Augmented SAMURAI for Long-Duration Drone Tracking

增强检测器的SAMURAI用于长时间无人机跟踪

Tamara R. Lenhard, Andreas Weinmann, Hichem Snoussi, Tobias Koch

机构 * Institute for the Protection of Terrestrial Infrastructures, German Aerospace Center (DLR)(地面基础设施保护研究所,德国航空航天中心(DLR)) ; ACIDA Lab, Technical University of Applied Sciences Würzburg-Schweinfurt(应用技术大学施魏尔堡-施维恩富特学院ACIDA实验室) ; Data Science Institute, European University of Technology(欧洲技术大学数据科学研究所) ; LIST3N, Université de Technologie de Troyes(图卢兹理工大学LIST3N)

AI总结 本文提出增强检测器的SAMURAI模型,用于提升无人机在复杂城市环境中的长时间跟踪鲁棒性,显著提高了成功率并降低了误检率。

Journal ref 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13007 2026-08-14 cs.CV 新提交

Structure-aware Riemannian Growth Fields for 4D Plant Modeling

面向4D植物建模的结构感知黎曼生长场

Meng-Yu Jennifer Kuo, Ryo Kawahara

机构 * Nara Women’s University(奈良女子大学) ; Kyoto University(京都大学)

AI总结 该研究提出结构感知黎曼生长场框架,用于从稀疏时间观测重建4D植物生长,构建了10天双物种标注数据集,在几何精度与对应一致性上优于现有方法。

Comments Accepted to WACV 2027 (Round 1)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17361 2026-08-12 cs.CV 版本更新

SuperQuadricOcc: Real-Time Self-Supervised Semantic Occupancy Estimation with Superquadric Volume Rendering

SuperQuadricOcc: 基于超级二次曲面体渲染的实时自监督语义占用估计

Seamie Hayes, Alexandre Boulch, Andrei Bursuc, Reenu Mohandas, Ganesh Sistu, Tim Brophy, Ciaran Eising

机构 * Data Driven Computer Engineering (D²iCE) Research Centre(数据驱动计算机工程(D²iCE)研究中心) ; University of Limerick(利默里克大学) ; Taighde Éireann – Research Ireland(爱尔兰研究——塔吉德)

AI总结 本文提出SuperQuadricOcc,通过超级二次曲面体渲染实现实时自监督语义占用估计,相比传统方法更高效且占用内存更少,在Occ3D-nuScenes数据集上取得最佳性能。

Comments Accepted at WACV 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09579 2026-08-11 cs.CV 新提交

You Only Flow Once: Calibrated and Real-Time Radar Pose Estimation with Multi-Hypothesis Normalizing Flows

仅一次流:基于多假设归一化流的校准且实时雷达位姿估计

Jonas Leo Mueller, Sebastian Hoefler, Dario Zanca, Naga Venkata Sai Jitin Jami, Thomas Altstidl, Bjoern M. Eskofier

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔兰根-纽伦堡大学) ; Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) ; LMU München(慕尼黑大学) ; Helmholtz Zentrum München(慕尼黑亥姆霍兹中心)

AI总结 该研究针对雷达位姿估计的歧义问题,提出MH-NFPG方法,结合时空Transformer与归一化流,实现高效实时的校准位姿估计,性能优于扩散模型。

Comments Accepted at the Winter Conference on Applications of Computer Vision (WACV) 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07559 2026-08-11 cs.CV cs.LG 新提交

MVMD: A Multi-View Approach for Enhanced Mirror Detection

MVMD:用于增强镜面检测的多视图方法

Yidan Shen, Yu Wen, Chen Zhang, Xin Fu, Renjie Hu

机构 * University of Houston(休斯顿大学)

AI总结 针对现有镜面检测仅关注单图像的局限,提出多视图镜面检测方法MVMD及首个多视图镜面检测数据库,通过三个模块提升检测效果,使准确率和IoU分别最高提升2.6%和11.1%,增强了镜面密集环境下的三维重建准确性。

Comments This work has already published at WACV 2025, just want more accessibility

Journal ref 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02927 2026-08-11 cs.CV cs.AI

PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding

PrismVAU: 用于多模态视频异常理解的提示优化推理系统

Iñaki Erregue, Kamal Nasrollahi, Sergio Escalera

机构 * Universitat de Barcelona(巴塞罗那大学) ; Computer Vision Center(计算机视觉中心) ; Aalborg University(奥胡斯大学) ; Milestone Systems(Milestone系统)

AI总结 PrismVAU通过轻量级系统和自动提示工程实现高效的多模态视频异常理解,无需复杂标注和外部模块。

Comments This paper has been accepted to the 6th Workshop on Real-World Surveillance: Applications and Challenges (WACV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04824 2026-08-07 cs.CV

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

SOVABench:多模态大语言模型的车辆监控动作检索基准

Oriol Rabasseda, Zenjie Li, Kamal Nasrollahi, Sergio Escalera

机构 * Milestone Systems A/S(Milestone Systems公司) ; Universitat de Barcelona(巴塞罗那大学) ; Computer Vision Center(计算机视觉中心) ; Aalborg Universitet(奥胡斯大学)

AI总结 SOVABench为多模态大语言模型提供车辆监控动作检索基准,通过定义两种评估协议评估跨动作区分和时间方向理解,展示了模型在复杂监控任务中的性能。

Comments This work has been accepted at Real World Surveillance: Applications and Challenges, 6th (in WACV Workshops)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14198 2026-08-07 eess.IV cs.CV 版本更新

Sparse Mixture-of-Experts for Non-Uniform Noise Reduction in MRI Images

用于MRI图像非均匀降噪的稀疏混合专家模型

Zeyun Deng, Joseph Campbell

机构 * Purdue University(普渡大学)

AI总结 针对MRI图像非均匀噪声问题,提出细粒度稀疏混合专家框架,将图像分区域后路由至专用降噪CNN,在合成与真实脑部MRI数据集上性能优于现有方法,且泛化性良好。

Comments Accepted to the WACV Workshop on Image Quality

Journal ref in 2025 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW), Tucson AZ USA, pp 260-268

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17109 2026-08-05 cs.CV

GMT: Guided Mask Transformer for Leaf Instance Segmentation

GMT:用于叶片实例分割的引导掩码Transformer

Feng Chen, Sotirios A. Tsaftaris, Mario Valerio Giuffrida

机构 * University of Edinburgh(爱丁堡大学) ; University of Nottingham(诺丁汉大学)

AI总结 针对叶片实例分割中叶片相似度高、遮挡多、标注数据少的难题,提出整合叶片空间分布先验的引导掩码Transformer GMT,在三个公开植物数据集上性能超越现有最优水平。

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2025 (Oral Presentation)

Journal ref IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2025, pp. 1217-1226

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02309 2026-08-04 cs.CV 新提交

CalibBEV: LiDAR-Camera Calibration via BEV Alignment

CalibBEV:基于鸟瞰图对齐的激光雷达-相机标定方法

Filippo D'Addeo, Lorenzo Cipelli, Adriano Cardace, Emanuele Ghelfi, Andrea Zinelli, Massimo Bertozzi

机构 * University of Bologna(博洛尼亚大学) ; University of Parma(帕尔马大学) ; Stanford University(斯坦福大学) ; VisLab srl(VisLab有限公司) ; Ambarella Inc.(安霸公司)

AI总结 CalibBEV是一种激光雷达-相机标定方法,通过两步BEV对齐结合CLIP对比损失实现跨模态统一,在KITTI、nuScenes基准上大幅降低了标定误差,达到最优性能。

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2026. p. 4345-4354

详情

展开后加载摘要…

URL PDF HTML 收藏
↑