arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

自动驾驶

自动驾驶感知、规划、BEV、占用预测、激光雷达和仿真评测。

共收录 21260 信号源:cs.RO, cs.CV, eess.IV, cs.AI

1. 感知 6072 篇

2504.08704 2026-01-27 cs.RO cs.LG 57%

Offline Reinforcement Learning using Human-Aligned Reward Labeling for Autonomous Emergency Braking in Occluded Pedestrian Crossing

利用人类对齐的奖励标注进行离线强化学习以实现自动驾驶中的遮挡行人横穿自主紧急制动

Vinal Asodia, Barkin Dagda, Yinglong He, Zhenhua Feng, Saber Fallah

机构 * CAV-Lab, School of Engineering(CAV实验室、工程学院) School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO

AI总结 本文提出了一种生成人类对齐奖励标签的方法,用于提升自动驾驶车辆在遮挡行人横穿场景中的安全性能。

Comments 39 pages, 14 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13373 2026-01-21 cs.CV 57%

A Lightweight Model-Driven 4D Radar Framework for Pervasive Human Detection in Harsh Conditions

一种轻量化的模型驱动4D雷达框架,用于在恶劣条件下进行 pervasive 人类检测

Zhenan Liu, Amir Khajepour, George Shaker

机构 * Mechanical \& Mechatronics Engineering University of Waterloo Waterloo, Canada Electrical \& Computer Engineering University of Waterloo Waterloo, Canada

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出了一种基于雷达的轻量模型驱动4D雷达框架,在恶劣环境中实现稳定的人体检测。

Journal ref IEEE PerCom 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13364 2026-01-21 cs.CV 57%

Real-Time 4D Radar Perception for Robust Human Detection in Harsh Enclosed Environments

实时4D雷达感知用于恶劣封闭环境中的可靠人类检测

Zhenan Liu, Yaodong Cui, Amir Khajepour, George Shaker

机构 * University of Waterloo(滑铁卢大学)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出了一种实时4D毫米波雷达感知方法,通过噪声过滤和基于规则的分类流程,在尘埃环境中实现可靠的人类检测。

Journal ref 2025 IEEE International Symposium on Antennas and Propagation and North American Radio Science Meeting (AP-S/CNC-USNC-URSI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13263 2026-01-21 cs.CV 57%

Deep Learning for Semantic Segmentation of 3D Ultrasound Data

用于3D超声数据语义分割的深度学习

Chenyu Liu, Marco Cecotti, Harikrishnan Vijayakumar, Patrick Robinson, James Barson, Mihai Caleap

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出了一种基于3D超声传感器的深度学习框架,用于实现3D语义分割,展示了其在恶劣环境下的稳健性能及潜在改进方向。

Comments 14 pages, 10 figures, 8 tables, presented at 2025 13th International Conference on Robot Intelligence Technology and Applications (RITA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05839 2026-01-21 cs.CV 57%

GeoSurDepth: Harnessing Foundation Model for Spatial Geometry Consistency-Oriented Self-Supervised Surround-View Depth Estimation

GeoSurDepth:利用基础模型实现以空间几何一致性为导向的自监督周围视图深度估计

Weimin Liu, Wenjun Wang, Joshua H. Meng

机构 * State Key Laboratory of Intelligent Green Vehicle and Mobility, School of Vehicle and Mobility, Tsinghua University, Beijing 100084, China(1 智能绿色车辆与移动国家重点实验室,车辆与移动学院,清华大学,北京100084,中国) California PATH, University of California, Berkeley, CA, United States(2 加州PATH,加州大学伯克利分校,加州,美国)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 GeoSurDepth通过利用基础模型和几何一致性,实现了更鲁棒的自监督周围视图深度估计,验证了其在自动驾驶中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10427 2026-01-21 cs.CV 57%

STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes

STRIDE-QA:用于城市驾驶场景时空推理的视觉问答数据集

Keishi Ishihara, Kento Sasaki, Tsubasa Takahashi, Daiki Shiono, Yu Yamaguchi

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 STRIDE-QA通过大规模视觉问答数据集提升自动驾驶中动态交通场景的时空推理能力,显著提升VLMs在空间定位和未来运动预测中的表现。

Comments Accepted to AAAI 2026 (Oral). project page: https://turingmotors.github.io/stride-qa/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18741 2026-01-21 eess.SY cs.AI cs.DC cs.LG cs.SY 57%

VREM-FL: Mobility-Aware Computation-Scheduling Co-Design for Vehicular Federated Learning

VREM-FL:面向车联网联邦学习的移动感知计算调度联合设计

Luca Ballotta, Nicolò Dal Fabbro, Giovanni Perin, Luca Schenato, Michele Rossi, Giuseppe Piro

专题命中 感知 :autonomous driving(abstract);分类 cs.AI

AI总结 VREM-FL通过结合车辆移动性和5G无线电环境图,优化车联网联邦学习的计算调度,提升模型训练效率和资源利用率。

Comments Copyright (c) 2024 IEEE. Personal use of this material is permitted. However, permission to use this material for any other purposes must be obtained from the IEEE by sending a request to pubs-permissions@ieee.org

Journal ref IEEE Transactions on Vehicular Technology, IEEE Transactions on Vehicular Technology, vol. 74, no. 2, pp. 3311-3326, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11779 2026-01-21 cs.CV 57%

Cross-Domain Object Detection Using Unsupervised Image Translation

跨领域目标检测使用无监督图像翻译

Vinicius F. Arruda, Rodrigo F. Berriel, Thiago M. Paixão, Claudine Badue, Alberto F. De Souza, Nicu Sebe, Thiago Oliveira-Santos

机构 * University of Trento (UNITN)(特伦托大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出了一种基于无监督图像翻译的方法,通过生成目标领域合成数据提升跨领域目标检测性能,实验表明在自动驾驶场景中优于现有方法。

Journal ref Expert Systems with Applications (ESWA), 192, 116334, 2022, Elsevier

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13883 2026-01-21 cs.CV 57%

YOLO-LLTS: Real-Time Low-Light Traffic Sign Detection via Prior-Guided Enhancement and Multibranch Feature Interaction

YOLO-LLTS: 通过先验引导增强和多分支特征交互实现实时低光交通标志检测

Ziyu Lin, Yunfan Wu, Yuhang Ma, Junzhou Chen, Ronghui Zhang, Jiaming Wu, Guodong Yin, Liang Lin

机构 * Guangdong Key Laboratory of Intelligent Transportation System, School of intelligent systems engineering, Sun Yat-sen University(广东智能交通系统重点实验室,智能系统工程学院,中山大学) Department of Architecture and Civil Engineering, Chalmers University of Technology(建筑与土木工程系,查尔姆斯理工大学) School of Mechanical Engineering, Southeast University(机械工程学院,东南大学) School of Computer Science and Engineering, Sun Yat-sen University(计算机科学与工程学院,中山大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 YOLO-LLTS通过先验引导增强和多分支特征交互提升低光环境下交通标志检测精度。

Comments This work has been published in IEEE Transactions on Instrumentation and Measurement

Journal ref IEEE Trans. Instrum. Meas., vol. 74, pp. 1-18, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10819 2026-01-19 cs.CV 57%

A Unified 3D Object Perception Framework for Real-Time Outside-In Multi-Camera Systems

面向实时多相机系统的统一3D物体感知框架

Yizhou Wang, Sameer Pusegaonkar, Yuxing Wang, Anqi Li, Vishal Kumar, Chetan Sethi, Ganapathy Aiyer, Yun He, Kartikay Thakkar, Swapnil Rathi, Bhushan Rupde, Zheng Tang, Sujit Biswas

机构 * NVIDIA Corporation(英伟达公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出了一种面向实时多相机系统的统一3D物体感知框架,通过优化的Sparse4D框架和生成式数据增强策略,实现了在大规模基础设施环境中的高效多目标跟踪与实时部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08355 2026-01-16 cs.CV 57%

Semantic Misalignment in Vision-Language Models under Perceptual Degradation

视觉-语言模型在感知退化下的语义错位

Guo Cheng

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本研究探讨了视觉-语言模型在感知退化下的语义错位问题,发现传统分割指标的下降并未影响下游行为,揭示了像素鲁棒性与多模态语义可靠性之间的脱节。

Comments 10 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22972 2026-01-16 cs.CV eess.SP 57%

Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection

基于小波的4D雷达张量与相机多视图融合用于鲁棒3D目标检测

Runwei Guan, Jianan Liu, Shaofeng Liang, Fangqiang Ding, Shanliang Yao, Xiaokai Bai, Daizong Liu, Tao Huang, Guoqiang Mao, Hui Xiong

机构 * Thrust of Artificial Intelligence, Hong Kong University of Science and Technology (Guangzhou)(人工智能 thrust,香港科技大学(广州)) Momoniai AI Department of Mechanical Engineering, Massachusetts Institute of Technology(机械工程系,麻省理工学院) School of Information Engineering, Yancheng Institute of Technology(信息工程学院,盐城科技学院) College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学) Institute for Math & AI, Wuhan University(数学与人工智能研究所,武汉大学) College of Science and Engineering and the Centre for AI and Data Science Innovation, James Cook University(科学与工程学院及人工智能与数据科学创新中心,詹姆斯库克大学) School of Transportation, Southeast University(交通运输学院,东南大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 WRCFormer通过小波注意力模块和几何引导渐进融合机制,高效融合4D雷达张量与相机图像,提升3D目标检测在恶劣天气下的鲁棒性。

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13601 2026-01-16 cs.CV 57%

Unleashing Semantic and Geometric Priors for 3D Scene Completion

释放语义和几何先验以实现3D场景补全

Shiyuan Chen, Wei Sui, Bohao Zhang, Zeyd Boukhers, John See, Cong Yang

机构 * D-Robotics

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 FoundationSSC通过双解耦机制和轴感知融合模块,提升3D场景补全的语义和几何指标表现。

Comments Accept by AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13165 2026-01-15 quant-ph cs.AI cs.LG 57%

QuFeX: Quantum feature extraction module for hybrid quantum-classical deep neural networks

QuFeX:用于混合量子-经典深度神经网络的量子特征提取模块

Naman Jain, Amir Kalev

机构 * Viterbi School of Engineering, University of Southern California, Los Angeles, California 90089, USA Information Sciences Institute, University of Southern California, Arlington, VA 22203, USA Department of Physics Center for Quantum Information Science \& Technology, University of Southern California, Los Angeles, California 90089, USA

专题命中 感知 :autonomous driving(abstract);分类 cs.AI

AI总结 QuFeX通过在混合量子-经典深度神经网络中引入量子特征提取模块,提升了图像分割任务的性能。

Comments V2: 17 pages, 15 figures, 2 Tables; published version

Journal ref Quantum Sci. Technol. 11 015017 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08042 2026-01-14 physics.app-ph cs.RO 57%

μDopplerTag: CNN-Based Drone Recognition via Cooperative Micro-Doppler Tagging

μDopplerTag: 基于协作微多普勒标记的CNN无人机识别

O. Yerushalimov, D. Vovchuk, A. Glam, P. Ginzburg

专题命中 感知 :LiDAR(abstract);分类 cs.RO

AI总结 本文提出基于电磁标签和CNN的无人机识别方法,利用微多普勒签名实现远距离高精度分类,适用于空域监控等关键应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12735 2026-01-14 cs.CV 57%

Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning

通过多模态提示调优对开放词汇目标检测器进行后门攻击

Ankita Raj, Chetan Arora

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 TrAP通过多模态提示调优对开放词汇目标检测器实施后门攻击,利用轻量级提示标记植入恶意行为,提升攻击成功率并改进下游任务性能。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07168 2026-01-14 cs.CV 57%

HisTrackMap: Global Vectorized High-Definition Map Construction via History Map Tracking

HisTrackMap: 通过历史地图追踪构建全局向量高精度地图

Jing Yang, Sen Yang, Xiao Tan, Hanli Wang

机构 * Tongji University(同济大学) Baidu Inc.(百度公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 HisTrackMap通过历史地图追踪构建全局向量高精度地图,提升时间连续性和几何构造质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22200 2026-01-13 cs.RO 57%

EnvoDat: A Large-Scale Multisensory Dataset for Robotic Spatial Awareness and Semantic Reasoning in Heterogeneous Environments

EnvoDat:一种大规模多感官数据集,用于机器人空间感知和异构环境中的语义推理

Linus Nwankwo, Bjoern Ellensohn, Vedant Dave, Peter Hofer, Jan Forstner, Marlene Villneuve, Robert Galler, Elmar Rueckert

机构 * Chair of Cyber-Physical System, Montanuniversität Leoben, Austria(智能物理系统系,莱布恩矿业大学,奥地利) Theresianische Militarakademie, Austria(特里西亚军事学院,奥地利) Chair of Subsurface Engineering, Montanuniversität Leoben, Austria(地下工程系,莱布恩矿业大学,奥地利)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO

AI总结 EnvoDat是一个大规模多感官数据集,用于提升机器人在复杂异构环境中的空间感知和语义推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19290 2026-01-13 cs.CV 57%

TRASE: Tracking-free 4D Segmentation and Editing

TRASE:无需跟踪的4D分割与编辑

Yun-Jin Li, Mariia Gladkova, Yan Xia, Daniel Cremers

机构 * TU Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 TRASE通过弱监督学习实现无需跟踪的4D分割,利用对比学习和聚类技术实现动态场景的高效分割与交互编辑。

Comments Accepted to 3DV 2026. Project page https://yunjinli.github.io/project-sadg

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04605 2026-01-09 cs.CV 57%

Detection of Deployment Operational Deviations for Safety and Security of AI-Enabled Human-Centric Cyber Physical Systems

面向AI赋能的人本型网络物理系统的部署操作偏差检测

Bernard Ngabonziza, Ayan Banerjee, Sandeep K. S. Gupta

专题命中 感知 :self-driving(abstract);分类 cs.CV

AI总结 本文提出了一种基于个性化图像的新技术,用于检测AI赋能的人本型网络物理系统在部署中的操作偏差,以确保其安全性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06282 2026-01-09 cs.CV 57%

From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot Learning

从数据集到现实世界:通过通用跨领域少样本学习实现通用3D目标检测

Shuangzhi Li, Junlong Shen, Lei Ma, Xingyu Li

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出通用跨领域少样本学习方法,通过融合2D语义与3D空间推理,实现对现实世界中常见和新类目标的高效检测。

Comments The latest version refines the few-shot setting on common classes, enforcing a stricter object-level definition

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03301 2026-01-08 cs.MA cs.AI 57%

PC2P: Multi-Agent Path Finding via Personalized-Enhanced Communication and Crowd Perception

PC2P:通过个性化增强通信与人群感知进行多智能体路径寻找

Guotao Li, Shaoyun Xu, Yuexing Hao, Yang Wang, Yuhui Sun

机构 * Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院微电子研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 感知 :occupancy(abstract);分类 cs.AI

AI总结 PC2P通过个性化增强通信与人群感知方法,提升多智能体路径寻找在复杂环境中的协同与扩展能力。

Comments 8 pages,7 figures,3 tables,Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01695 2026-01-06 cs.CV 57%

Learnability-Driven Submodular Optimization for Active Roadside 3D Detection

基于可学习性的子模优化用于主动道路3D检测

Ruiyu Mao, Baoming Zhang, Nicholas Ruozzi, Yunhui Guo

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本研究提出了一种基于可学习性的主动学习框架,用于道路侧单目3D物体检测,通过选择信息丰富且可可靠标注的场景,有效减少标注成本并提升模型性能。

Comments 10 pages, 7 figures. Submitted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01676 2026-01-06 cs.CV 57%

LabelAny3D: Label Any Object 3D in the Wild

LabelAny3D: 在真实世界中标注任意物体的3D

Jin Yao, Radowan Mahmud Redoy, Sebastian Elbaum, Matthew B. Dwyer, Zezhou Cheng

机构 * University of Virginia(弗吉尼亚大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 LabelAny3D通过分析-合成框架生成高质量3D标注,提升单目3D检测性能,推动真实世界3D识别发展。

Comments NeurIPS 2025. Project page: https://uva-computer-vision-lab.github.io/LabelAny3D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00368 2026-01-05 cs.CV 57%

Mask-Conditioned Voxel Diffusion for Joint Geometry and Color Inpainting

基于掩码的体素扩散用于联合几何和颜色修复

Aarya Sumuk

专题命中 感知 :occupancy(abstract);分类 cs.CV

AI总结 本文提出基于掩码的体素扩散方法,用于修复受损3D物体的几何和颜色,通过两阶段框架实现更完整和一致的修复效果。

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24680 2026-01-01 cs.RO 57%

ReSPIRe: Informative and Reusable Belief Tree Search for Robot Probabilistic Search and Tracking in Unknown Environments

ReSPIRe: 信息性和可重用的信念树搜索用于未知环境中的机器人概率搜索与跟踪

Kangjie Zhou, Zhaoyang Li, Han Gao, Yao Su, Hangxin Liu, Junzhi Yu, Chang Liu

专题命中 感知 :trajectory planning(abstract);分类 cs.RO

AI总结 ReSPIRe通过分层粒子结构和可重用信念树搜索,在未知环境中实现高效且稳定的机器人目标搜索与跟踪。

Comments 17 pages, 12 figures, accepted to IEEE Transactions on Systems, Man, and Cybernetics: Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24243 2026-01-01 cs.CV 57%

MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic Segmentation

MambaSeg: 利用Mamba实现准确且高效的图像事件语义分割

Fuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji, Chao Chen, Qingyi Gu, Zhenliang Ni

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 MambaSeg通过双分支框架和双维交互模块,实现高效准确的多模态语义分割,适用于快速运动和低光条件下的图像事件分割任务。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24227 2026-01-01 cs.CV 57%

Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes

Mirage: 驾驶场景中光实且一致的资产编辑一步视频扩散

Shuyun Wang, Haiyang Sun, Bing Wang, Hangjun Ye, Xin Yu

机构 * The University of Queensland(昆士兰大学) Xiaomi EV(小米电动车)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 Mirage提出了一种用于驾驶场景中光实且一致的资产编辑的一步视频扩散模型,通过引入时序无关的潜在特征和两阶段数据对齐策略,提升了视觉保真度和时间一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23635 2025-12-30 cs.CV 57%

Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception

重新思考端到端3D感知的时空对齐

Xiaoyu Li, Peidong Li, Xian Wu, Long Shi, Dedong Liu, Yitao Wu, Jiajia Fu, Dixiao Cui, Lijun Zhao, Lining Sun

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 HAT通过自适应解码多假设生成最优时空对齐提案,提升自动驾驶中3D感知精度和鲁棒性。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23585 2025-12-30 cs.RO cs.SY eess.SY 57%

Unsupervised Learning for Detection of Rare Driving Scenarios

无监督学习用于罕见驾驶场景检测

Dat Le, Thomas Manhardt, Moritz Venator, Johannes Betz

机构 * Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University Munich(自主车辆系统教授职位,技术大学慕尼黑工程与设计学院) Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所) CARIAD SE(CARIAD公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO

AI总结 本研究提出一种无监督学习方法,利用深度隔离森林和t-SNE技术,有效检测自动驾驶中的罕见危险场景。

详情

展开后加载摘要…

URL PDF HTML 收藏