arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Winter Conference on Applications of Computer Vision · 会议 · Computer Vision

2025-12-19 至 2025-12-19 共收录 4
2512.16818 2025-12-19 cs.CV

DenseBEV: Transforming BEV Grid Cells into 3D Objects

DenseBEV:将BEV网格单元转换为3D对象

Marius Dähling, Sebastian Krebs, J. Marius Zöllner

机构 * NMS Non-Maximum Suppression IoU Intersection-over-Union FoV Field of View mAP mean Average Precision NDS nuScenes detection score ATE Average Translation Error ASE Average Scale Error AOE Average Orientation Error AVE Average Velocity Error AAE Average Attribute Error LET Longitudinal Error Tolerant APL Average Precision Longitudinal APH Average Precision Heading LSS Lift, Splat, Shoot BEV Bird’s-Eye-View DenseBEV: Transforming BEV Grid Cells into 3D Objects(NMS非最大抑制IoU交并比FoV视野mAP均值平均精度NDSnuScenes检测得分ATE平均平移误差ASE平均尺度误差AOE平均方位误差AVE平均速度误差AAE平均属性误差LET纵向误差容忍APL纵向平均精度APH平均方位精度LSS抬升、投射、射击BEV鸟视图DenseBEV:将BEV网格单元转换为3D对象)

AI总结 DenseBEV通过直接使用BEV特征单元作为锚点,提出了一种高效且直观的多摄像头3D目标检测方法,提升了小目标检测性能。

Comments 15 pages, 8 figures, accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16133 2025-12-19 cs.CV

Interaction-via-Actions: Cattle Interaction Detection with Joint Learning of Action-Interaction Latent Space

通过动作交互:基于联合学习动作-交互潜在空间的牛群交互检测

Ren Nakagawa, Yang Yang, Risa Shinoda, Hiroaki Santo, Kenji Oyama, Fumio Okura, Takenao Ohkawa

机构 * Kobe University(神户大学) The University of Osaka(大阪大学)

AI总结 本文提出CattleAct方法,通过联合学习动作-交互潜在空间,实现高效检测牛群交互,提升智能畜牧业管理能力。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15957 2025-12-19 cs.CV cs.AI

Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models

看见即信仰(并预测):基于视觉语言模型的上下文感知多人类行为预测

Utsav Panchal, Yuchen Liu, Luigi Palmieri, Ilche Georgievski, Marco Aiello

机构 * Institute of Architecture of Application Systems, University of Stuttgart, Germany(应用系统建筑研究所,斯图加特大学,德国) Bosch Research, Germany(博世研究,德国)

AI总结 CAMP-VLM通过结合视觉语言模型与上下文特征,提升了多人类行为预测的准确性,其在预测精度上比基线模型高66.9%。

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15949 2025-12-19 cs.CV

The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs

感知观测站:多模态大语言模型的鲁棒性与基础性表征

Tejas Anvekar, Fenil Bardoliya, Pavan K. Turaga, Chitta Baral, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学)

AI总结 感知观测站通过系统性扰动和真实数据集,评估多模态大语言模型在视觉基础性和关系结构上的鲁棒性,揭示其在扰动下的表现,为模型分析提供系统基础。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏