arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

自动驾驶

自动驾驶感知、规划、BEV、占用预测、激光雷达和仿真评测。

共收录 21273 信号源:cs.RO, cs.CV, eess.IV, cs.AI

1. 感知 6072 篇

2304.02991 2023-04-07 cs.CV 61%

Exploiting the Complementarity of 2D and 3D Networks to Address Domain-Shift in 3D Semantic Segmentation

Adriano Cardace, Pierluigi Zama Ramirez, Samuele Salti, Luigi Di Stefano

专题命中 感知 :autonomous driving(abstract,comments);分类 cs.CV

Comments Accepted at the CVPR2023 Workshop on Autonomous Driving (WAD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10692 2022-11-10 cs.CV cs.LG 61%

Multi-level Domain Adaptation for Lane Detection

Chenguang Li, Boheng Zhang, Jia Shi, Guangliang Cheng

专题命中 感知 :autonomous driving(abstract,comments);分类 cs.CV

Comments Proceedings of the CVPR 2022 Workshop of Autonomous Driving

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16204 2022-10-31 cs.CV 61%

TripletTrack: 3D Object Tracking using Triplet Embeddings and LSTM

Nicola Marinello, Marc Proesmans, Luc Van Gool

专题命中 感知 :autonomous driving(abstract,comments);分类 cs.CV

Comments Accepted to CVPR 2022 Workshop on Autonomous Driving

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops June 2022 4500-4510

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15074 2021-10-29 cs.CV 61%

Meta Guided Metric Learner for Overcoming Class Confusion in Few-Shot Road Object Detection

Anay Majee, Anbumani Subramanian, Kshitij Agrawal

专题命中 感知 :autonomous driving(abstract,comments);分类 cs.CV

Comments Accepted to NeurIPS 2021 Workshop on Machine Learning For Autonomous Driving, 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.11422 2021-06-23 cs.CV cs.LG 61%

MODETR: Moving Object Detection with Transformers

Eslam Mohamed, Ahmad El-Sallab

专题命中 感知 :autonomous driving(abstract,journal_ref);分类 cs.CV

Journal ref Machine Learning for Autonomous Driving Workshop at the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.11859 2020-09-25 cs.CV cs.LG 61%

Multi-Frame to Single-Frame: Knowledge Distillation for 3D Object Detection

Yue Wang, Alireza Fathi, Jiajun Wu, Thomas Funkhouser, Justin Solomon

专题命中 感知 :autonomous driving(abstract,comments);分类 cs.CV

Comments The Workshop on Perception for Autonomous Driving at ECCV2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.11757 2019-09-02 cs.CV cs.LG 61%

Temporal Coherence for Active Learning in Videos

Javad Zolfaghari Bengar, Abel Gonzalez-Garcia, Gabriel Villalonga, Bogdan Raducanu, Hamed H. Aghdam, Mikhail Mozerov, Antonio M. Lopez, Joost van de Weijer

专题命中 感知 :autonomous driving(abstract,comments);分类 cs.CV

Comments Accepted at ICCVW 2019 (CVRSUAD-Road Scene Understanding and Autonomous Driving)

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.08492 2019-04-23 cs.CV 61%

MultiNet++: Multi-Stream Feature Aggregation and Geometric Loss Strategy for Multi-Task Learning

Sumanth Chennupati, Ganesh Sistu, Senthil Yogamani, Samir A Rawashdeh

专题命中 感知 :autonomous driving(abstract,comments);分类 cs.CV

Comments Accepted for CVPR 2019 Workshop on Autonomous Driving (WAD). Demo Video can be accessed at https://youtu.be/E378PzLq7lQ

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.02503 2019-03-07 cs.RO 61%

The AI Driving Olympics at NeurIPS 2018

Julian Zilly, Jacopo Tani, Breandan Considine, Bhairav Mehta, Andrea F. Daniele, Manfred Diaz, Gianmarco Bernasconi, Claudio Ruch, Jan Hakenberg, Florian Golemo, A. Kirsten Bowser, Matthew R. Walter, Ruslan Hristov, Sunil Mallya, Emilio Frazzoli, Andrea Censi, Liam Paull

专题命中 感知 :autonomous driving(abstract);分类 cs.RO;self-driving(comments)

Comments Competition, robotics, safety-critical AI, self-driving cars, autonomous mobility on demand, Duckietown

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.05998 2018-12-11 cs.CV 61%

Minimizing Supervision for Free-space Segmentation

Satoshi Tsutsui, Tommi Kerola, Shunta Saito, David J. Crandall

专题命中 感知 :autonomous driving(abstract,comments);分类 cs.CV

Comments Link to source code added; Typo fixed from the version published in CVPR 2018 Workshop on Autonomous Driving (WAD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09305 2026-08-25 cs.CV 版本更新 57%

VAGNet: Vision-based Accident Anticipation with Global Features

VAGNet:基于视觉的事故预测与全局特征

Vipooshan Vipulananthan, Charith D. Chitraranjan

机构 * Department of Computer Science and Engineering, University of Moratuwa(moratuwa大学计算机科学与工程系)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出VAGNet,通过全局特征预测事故,提升实时性与效率,实验显示在四个数据集上精度和响应时间均优于现有方法。

Comments Published in IEEE Open Journal of Vehicular Technology (OJVT)

Journal ref IEEE Open Journal of Vehicular Technology, vol. 7, pp. 1873-1883, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15052 2026-08-25 cs.RO 版本更新 57%

CAVERS: Multimodal SLAM Data from a Natural Karstic Cave with Ground Truth Motion Capture

CAVERS:从自然喀斯特洞穴获取多模态SLAM数据并采用地面真实运动捕捉

Giacomo Franchini, David Rodríguez-Martínez, Alfonso Martínez-Petersen, C. J. Pérez-del-Pulgar, Marcello Chiaberge

机构 * Polytechnic of Turin Interdepartmental Centre for Service Robotics (PIC4SeR)(都灵理工大学服务机器人跨部门研究中心(PIC4SeR)) Systems Engineering and Automation Department, Universidad de Málaga(马德里大学系统工程与自动化系)

专题命中 感知 :LiDAR(abstract);分类 cs.RO

AI总结 本文提出CAVERS数据集,通过地面真实运动捕捉系统提供高精度姿态和速度数据,验证了多模态传感器在复杂洞穴环境中的SLAM算法性能。

Comments 8 pages, 4 figures, accepted version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12423 2026-08-25 cs.CV math.OC 版本更新 57%

HOT-POT: Optimal Transport for Sparse Stereo Matching

HOT-POT:稀疏立体匹配中的最优传输

Antonin Clerc, Michael Quellmalz, Moritz Piening, Philipp Flotho, Gregor Kornhardt, Gabriele Steidl

机构 * Univ. Bordeaux, CNRS, Bordeaux INP, IMB, UMR 5251(波尔多大学,CNRS,波尔多INP,IMB,UMR 5251) Okinawa Institute of Science and Technology(冲绳科学技术研究所)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 HOT-POT通过最优传输方法解决稀疏立体匹配问题,利用epipolar距离和3D射线距离提升匹配效率,应用于面部分析中的地标匹配。

Comments 14 pages, 9 figures, 8 tables

Journal ref Journal of Mathematical Imaging and Vision 68(6), pages 951-976, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19804 2026-08-21 cs.AI 新提交 57%

ADAPT: Physics-Aware Diffusion-based World Models for Adaptive Predictive Transferable HVAC Control

ADAPT:面向自适应可迁移HVAC控制的物理感知扩散式世界模型

Xu Yang, Kailai Sun, Dianyu Zhong, Qianchuan Zhao

专题命中 感知 :occupancy(abstract);分类 cs.AI

AI总结 本文针对现有HVAC控制方法泛化性差的问题,提出物理感知扩散世界模型ADAPT,在IID控制下可降HVAC能耗7.3%、不适度30.2%,OOD场景下迁移鲁棒性显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19380 2026-08-21 cs.CV 新提交 57%

CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios

CAViAR:用于真实场景细粒度事故推理的因果视频数据集

Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G, Abhishek Aich

机构 * NEC Laboratories, America(美国NEC实验室)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 该研究推出人工标注的真实事故视频基准CAViAR,测试发现现有VLMs存在感知-推理差距,无法可靠将驾驶场景主体行为映射到责任类别。

Comments Accepted to ECCV 2026 Workshop DriveX

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19298 2026-08-21 cs.CV 新提交 57%

SceneGTMM: A Conformal Mapping-based Scene-Aware Transferable GNN-Transformer Dual-Graph Interaction Framework for Map Matching

SceneGTMM:一种基于保角映射的场景感知可迁移GNN-Transformer双图交互地图匹配框架

Yongliang Zhang, Feng Song, Ji Chen, Lishuai Guo, Yong Deng, Yue Zheng, Tianyi Liu, Zhixiong Chen, Qixin Zhang

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出SceneGTMM框架,通过保角映射场景策略、双图交互架构与CRF增强预测,提升地图匹配的噪声鲁棒性、跨区域迁移性与可解释性,在多源及跨城轨迹匹配中表现优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19188 2026-08-20 cs.RO 新提交 57%

PartialBiGrasp: Inferring Hidden Local Geometry for Bimanual Grasping from Partial Views

PartialBiGrasp:从部分视角推断隐藏局部几何以实现双臂抓取

Ayush Kaura, Vignesh Vembar, Md Faizal Karim, Keshab Patra, K Madhava Krishna

机构 * Robotics Research Center IIIT Hyderabad(印度信息技术研究所海得拉巴分校机器人研究中心)

专题命中 感知 :occupancy(abstract);分类 cs.RO

AI总结 PartialBiGrasp是基于部分点云的双臂抓取生成框架,通过卷积占用网络学习几何特征生成抓取对并优化,经实验验证可实现鲁棒稳定的抓取。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19036 2026-08-20 cs.CV 新提交 57%

USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes

USR-Drive:通过3D高斯与边界框联合去噪实现统一驾驶场景表示

Li-Heng Chen, Haokai Pang, Chengye Su, Jiarun Liu, Qifeng Chen, Ziqian Ni, Jianxin Huang, Shi-Sheng Huang, Hongbo Fu, Sheng Yang

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本文提出USR-Drive框架,将3D高斯与边界框作为对齐的潜在令牌流,通过统一多模态扩散Transformer联合去噪,在nuScenes和VKitti数据集上实现动态重建与3D检测的SOTA性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15048 2026-08-20 cs.CV 版本更新 57%

RoGS: Adaptive Meshgrid Gaussian for Large-Scale Road Surface Mapping

RoGS:用于大规模路面映射的自适应网格高斯方法

Tianchen Deng, Zhiheng Feng, Wenhua Wu, Ziming Li, Chang Nie, Siting Zhu, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong university(上海交通大学自动化与智能感知学院) State Key Laboratory of Avionics Integration and Aviation System-of-Systems Synthesis(航空电子集成与航空系统-of-系统综合国家重点实验室) Shanghai Key Laboratory of Navigation and Location Based Services(上海基于位置服务重点实验室)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 针对大规模路面映射中传统方法的局限,提出ROADGS-T框架,基于自适应网格高斯表示,通过放置二维高斯面片建模路面,减少冗余,并引入自适应网格和姿态稳健细化策略,提升表示效率和结构保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15796 2026-08-18 cs.CV 新提交 57%

Emergent 3D Instance Segmentation from Self-Supervised Point Transformers

自监督点Transformer的涌现式3D实例分割

Ted Lentsch, Santiago Montiel-Marín, Holger Caesar, Julian F. P. Kooij

机构 * Delft University of Technology(代尔夫特理工大学) University of Alcalá(阿尔卡拉大学)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本研究探究自监督点Transformer是否含分离3D实例的结构信息,据此提出无需训练的TokenGraph3D方法,在无先验条件下大幅优于基线,使涌现的3D实例结构可见。

Comments ECCV 2026 DriveX

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12122 2026-08-17 cs.RO 版本更新 57%

OctoSplat: Hybrid OctoMap-Gaussian Splatting for Active Semantic Mapping and Phenotyping with Horticultural Robots

利用高斯点云进行园艺环境的主动语义映射

Jose Cuaran, Naveen K. Uppalapati, Girish Chowdhary

机构 * the Siebel School of Computing and Data Science(塞比尔计算与数据科学学院) the Department of Agricultural and Biological Engineering(农业与生物工程系) National Center for Supercomputing Applications at University of Illinois, Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校国家超级计算中心)

专题命中 感知 :occupancy(abstract);分类 cs.RO

AI总结 本文提出利用移动机械臂和高斯点云进行园艺环境主动3D重建,提升重建精度与效率,实现水果计数和体积估算。

Comments 23 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12608 2026-08-14 cs.CV 版本更新 57%

A Data Efficiency Study of Synthetic Fog for Object Detection Using the Clear2Fog Pipeline

使用Clear2Fog管道研究合成雾的数据效率

Mohamed Ahmed Mohamed, Xiaowei Huang

机构 * Waymo Open Dataset(Waymo开放数据集)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 本文提出Clear2Fog管道,通过物理模拟生成雾数据,研究环境多样性对模型鲁棒性的影响,发现混合密度数据训练模型优于固定密度数据,且通过调整学习率可克服合成偏见。

Comments Accepted to Neurocomputing. Project code and experimental configs available at https://github.com/mmohamed28/Clear2Fog

Journal ref Neurocomputing, Vol. 703, 134798, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05474 2026-08-13 cs.CV 版本更新 57%

3D Scene Generation: A Survey

3D场景生成:一项综述

Haozhe Xie, Beichen Wen, Zhaoxi Chen, Fangzhou Hong, Ziwei Liu

机构 * S-Lab, Nanyang Technological University, Singapore(南洋理工大学S实验室)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 本综述系统梳理3D场景生成的四大范式方法,分析其技术基础与挑战,展望物理感知生成等方向,为该领域发展提供参考。

Comments Accepted by IJCV. Project Page: https://github.com/hzxie/Awesome-3D-Scene-Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03314 2026-08-12 cs.CV 版本更新 57%

TASE: Truncation-Aware Semantic Embeddings for 3D Scene Understanding and Editing

TASE: 用于3D场景理解与编辑的截断感知语义嵌入

Tim-Felix Faasch, Jochen Kall, Lucas Nunes, Jens Behley, Cyrill Stachniss

机构 * Bosch Research(博世研究院) Rheinisch-Westfälische Technische Hochschule Aachen(亚琛工业大学) University of Bonn(波恩大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 提出TASE方法,通过将预训练的2D语义特征投影到截断感知嵌入空间,结合尺度和平移等变损失,实现可控的3D场景文本驱动编辑,在大几何修改任务上显著优于现有方法。

Comments Code: https://github.com/boschresearch/TASE

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01656 2026-08-12 cs.CV 版本更新 57%

WaveInst: A Frequency-Domain Enhanced Network for Fine-Grained Thin Tree Trunk Extraction in Forest Scenes

WaveInst:面向森林场景中细粒度细树干提取的频域增强网络

Chenyang Fan, Xujie Zhu, Taige Luo, Zhulin Chen, Sheng Xu

机构 * College of Information Science and Technology & Artificial Intelligence, Nanjing Forestry University(南京林业大学信息科学与人工智能学院) Department of Geography, College of Natural Resources and Environment, Virginia Tech(维吉尼亚理工大学地理系) Institute of Forest Resource Information Techniques, Chinese Academy of Forestry(中国林业科学研究院森林资源信息技术研究所) State Forestry and Grassland Administration, Key Laboratory of Forest Management and Growth Modelling(国家林业和草原局森林管理与生长建模重点实验室) School of Geospatial Artificial Intelligence, East China Normal University(华东师范大学地理空间人工智能学院)

专题命中 感知 :LiDAR(abstract);分类 cs.CV

AI总结 针对现有方法在森林场景细树干提取中易误判、受数据与直径差异限制的问题,提出WaveInst频域增强实例分割网络,在多数据集上实现了更优的细树干提取性能,尤其在幼龄树干上优势显著。

Comments 18 pages, 16 figures, published in IEEE Transactions on Geoscience and Remote Sensing

Journal ref IEEE Transactions on Geoscience and Remote Sensing, vol. 64, pp. 4409718-4409718, 2026, Art no. 4409718

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09597 2026-08-11 cs.CV 新提交 57%

ResemBrick: Brick Reconstruction from Photographs with Perceptual Fidelity and Buildability

ResemBrick:从照片重建兼具感知保真度与可建造性的积木模型

Xilun Chen, Hanwen Wan, Yusong Zhao, Zexin Lin, Ruixiang Liao, Xiaoqiang Ji

专题命中 感知 :occupancy(abstract);分类 cs.CV

AI总结 ResemBrick将离散化与组装耦合,通过预算占用补全和可建造性构建,生成兼具感知保真度与可建造性的积木模型,性能优于现有系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09493 2026-08-11 cs.CV 新提交 57%

GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

GeoRoute:面向交通未来帧预测的几何感知混合推理方法

Khang Minh Le, Hieu Dinh Trung Pham, Luu Thanh Danh, Nam-Tien Le, Hieu Anh Ngo, Phuong Huu Vu Tran, Son Nguyen Minh Le, Nguyen Trong Nghia, Tu Tran Thi Cam, Huy Minh Nhat Nguyen, Cuong Tuan Nguyen

机构 * Vietnamese-German University(越南-德国大学) University of Science, Ho Chi Minh City(胡志明市科学大学) Ho Chi Minh City University of Technology(胡志明市技术大学) University of Information Technology, Ho Chi Minh City(胡志明市信息技术大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 GeoRoute是一种无需训练的几何感知混合推理框架,通过多帧深度分层渲染器和视角条件选择预测器,在AI City Challenge Track 5基准上实现了具有竞争力的交通未来帧预测性能,提升了静态几何稳定性。

Comments accepted to the ECCV 2026 AI City Challenge Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08947 2026-08-11 cs.CV cs.HC cs.LG 新提交 57%

Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis

网络摄像头眼动能否约束驾驶模型中的 mesa 目标?一项仪器精度分析

Lennox Anderson, Ahmed Boutar, Jonah Mulcrone, Tal Erez

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 该研究探究能否用 WebGazer 眼动数据约束自动驾驶模型的 mesa 目标,经多组实验发现无统计显著效果,原因是 WebGazer 误差过大无法实现物体级注视归因。

Comments 6 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08476 2026-08-11 cs.CV 新提交 57%

RayLift: Lifting Complementary Ray-Wise Evidence with 3D Geometry Priors for Semantic Scene Completion

RayLift:利用3D几何先验提升互补射线级证据以实现语义场景补全

Meng Wang, Hongxia Yu, Wenzhe He, Xingdong Song, Huilong Pi, Jiapeng Zhang, Ruihui Li

专题命中 感知 :autonomous driving(abstract);分类 cs.CV

AI总结 针对现有语义场景补全方法的深度误差传播问题,提出RayLift框架,通过互补上下文编码器、深度射线证据提升模块和语义感知体素集成器,在两个基准数据集上实现优于现有方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07643 2026-08-11 cs.CV cs.LG 新提交 57%

Data collection from highways: a geometric, class-agnostic approach to embedded vehicle counting

高速公路数据采集:一种用于嵌入式车辆计数的几何、类别无关方法

Lucas Gouveia Omena Lopes, William W. M. Lira, Alexandre M. Lima, Thales M. A. Vieira

机构 * Universidade Federal de Alagoas(阿拉戈斯联邦大学)

专题命中 感知 :occupancy(abstract);分类 cs.CV

AI总结 本文提出一种面向嵌入式设备的类别无关几何车辆计数方法,无需物体模型与训练集,在树莓派硬件上快于实时运行,实地部署准确率达91%,适用于无标注数据等场景。

详情

展开后加载摘要…

URL PDF HTML 收藏