arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

自动驾驶

自动驾驶感知、规划、BEV、占用预测、激光雷达和仿真评测。

共收录 731 信号源:cs.RO, cs.CV, eess.IV, cs.AI

1. 端到端驾驶 731 篇

2602.03213 2026-02-11 cs.CV 57%

ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask

ConsisDrive: 用于视频生成的实例掩码驾驶世界模型

Zhuoran Yang, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.CV

AI总结 ConsisDrive通过实例掩码注意力和损失机制提升驾驶视频生成质量及自动驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03242 2026-02-04 cs.CV 57%

InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation

InstaDrive: 为真实和一致的视频生成实例感知的驾驶世界模型

Zhuoran Yang, Xi Guo, Chenjing Ding, Chiyu Wang, Wei Wu, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) SenseAuto(感etime) Tsinghua University(清华大学)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.CV

AI总结 InstaDrive通过实例感知机制提升驾驶视频生成质量,增强自动驾驶任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02864 2026-02-04 cs.RO 57%

Accelerating Structured Chain-of-Thought in Autonomous Vehicles

加速自动驾驶中的结构化链式推理

Yi Gu, Yan Wang, Yuxiao Chen, Yurong You, Wenjie Luo, Yue Wang, Wenhao Ding, Boyi Li, Heng Yang, Boris Ivanovic, Marco Pavone

机构 * NVIDIA University of Southern California(南加州大学) Harvard University(哈佛大学) Stanford University(斯坦福大学)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.RO

AI总结 FastDriveCoT通过并行解码方法加速自动驾驶中的结构化链式推理,实现3-4倍的生成速度提升和端到端延迟的显著降低,同时保持推理效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12432 2026-01-21 cs.CV cs.MM 57%

SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition

SkeFi: 无线传感器跨模态知识转移用于基于骨骼的动作识别

Shunyu Huang, Yunjiao Zhou, Jianfei Yang

机构 * School of Electrical and Electronics Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院)

专题命中 端到端驾驶 :LiDAR(abstract);分类 cs.CV

AI总结 SkeFi通过跨模态知识转移方法,利用无线传感器提升基于骨骼的动作识别性能,实现毫米波和LiDAR上的先进表现。

Comments Published in IEEE Internet of Things Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09452 2026-01-15 cs.CV 57%

MAD: Motion Appearance Decoupling for efficient Driving World Models

MAD:用于高效驾驶世界模型的运动外观解耦

Ahmad Rahimi, Valentin Gerard, Eloi Zablocki, Matthieu Cord, Alexandre Alahi

机构 * Sorbonne Université(索邦大学)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.CV

AI总结 MAD通过解耦运动学习与外观合成,高效地将通用视频扩散模型转化为可控的驾驶世界模型,实现低计算成本和高性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06223 2026-01-13 cs.CY cs.AI 57%

Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness

迈向安全和负责任的AI代理:透明、问责和可信的三支柱模型

Edward C. Cheng, Jeshua Cheng, Alice Siu

机构 * InquiryOn

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.AI

AI总结 本文提出三支柱模型,旨在通过透明、问责和可信机制,确保AI代理的安全性和责任性,促进负责任的AI发展。

Comments 15 pages, 8 figures, conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02228 2026-01-06 cs.LG cs.AI 57%

Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning

世界模型在线模仿学习中的耦合分布随机专家蒸馏

Shangzhe Li, Zhiao Huang, Hao Su

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Hillbot(Hillbot公司) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.AI

AI总结 本研究提出一种基于随机网络蒸馏的耦合分布随机专家蒸馏方法,用于提升世界模型在线模仿学习的稳定性与性能。

Comments NeurIPS 2025 Workshop of Embodied World Models; Code Available at: https://github.com/TobyLeelsz/CDRED-WM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14499 2025-12-11 cs.CL cs.AI cs.LG 57%

Understanding World or Predicting Future? A Comprehensive Survey of World Models

理解世界还是预测未来?世界模型的全面综述

Jingtao Ding, Yunke Zhang, Yu Shang, Jie Feng, Yuheng Zhang, Zefang Zong, Yuan Yuan, Hongyuan Su, Nian Li, Jinghua Piao, Yucheng Deng, Nicholas Sukiennik, Chen Gao, Fengli Xu, Yong Li

机构 * Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University(电子工程系,信息科学与技术国家研究中心(BNRist),清华大学)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.AI

AI总结 本文综述了世界模型在理解与预测方面的功能,探讨了其在多个领域的应用及未来研究方向。

Comments Extended version of the original ACM CSUR paper, 49 pages, 6 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01119 2025-12-02 cs.LG cs.AI 57%

World Model Robustness via Surprise Recognition

通过惊喜识别提升世界模型鲁棒性

Geigh Zollicoffer, Tanush Chopra, Mingkuan Yan, Xiaoxu Ma, Kenneth Eaton, Mark Riedl

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 端到端驾驶 :self-driving(abstract);分类 cs.AI

AI总结 通过惊喜识别提升世界模型在噪声环境下的鲁棒性,增强自动驾驶模拟中智能体的稳定性与性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00005 2025-12-02 cs.RO 57%

DREAMer-VXS: A Latent World Model for Sample-Efficient AGV Exploration in Stochastic, Unobserved Environments

DREAMer-VXS:一种用于在随机、未观测环境中高效样本AGV探索的潜在世界模型

Agniprabha Chakraborty

机构 * Department of Power Engineering, Jadavpur University(功率工程系,贾瓦德普尔大学)

专题命中 端到端驾驶 :LiDAR(abstract);分类 cs.RO

AI总结 DREAMer-VXS通过高效样本利用和潜在世界模型,实现AGV在随机未观测环境中的高效探索与鲁棒导航

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11393 2025-11-17 cs.AI 57%

Robust and Efficient Communication in Multi-Agent Reinforcement Learning

Zejiao Liu, Yi Li, Jiali Wang, Junqi Tu, Yitian Hong, Fangfei Li, Yang Liu, Toshiharu Sugawara, Yang Tang

机构 * Department of Mathematics, East China University of Science and Technology(东华大学数学系) Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, East China University of Science and Technology(能源化工过程智能制造重点实验室,东华大学) Hangzhou School of Automation, Zhejiang Normal University(浙江师范大学杭州自动化学院) Department of Computer Science, Waseda University(早稻田大学计算机科学系)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06071 2025-11-12 cs.CV 57%

TransParking: A Dual-Decoder Transformer Framework with Soft Localization for End-to-End Automatic Parking

Hangyu Du, Chee-Meng Chew

机构 * College of Design and Engineering, National University of Singapore(设计与工程学院,新加坡国立大学)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02851 2025-11-07 cs.RO cs.DC 57%

Action Deviation-Aware Inference for Low-Latency Wireless Robots

Jeyoung Park, Yeonsub Lim, Seungeun Oh, Jihong Park, Jinho Choi, Seong-Lyun Kim

机构 * Department of MME, University of Waterloo(水坝大学机械工程系) School of EEE, Yonsei University(延世大学电子工程系) ISTD Pillar, Singapore University of Technology and Design(新加坡技术与设计大学ISTD支柱) School of EME, University of Adelaide(阿德莱德大学电子工程系)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07416 2025-11-04 cs.LG cs.AI 57%

LiteVLM: A Low-Latency Vision-Language Model Inference Pipeline for Resource-Constrained Environments

Jin Huang, Yuchao Jin, Le An, Josh Park

机构 * NVIDIA(英伟达)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19578 2025-10-23 cs.CV 57%

VGD: Visual Geometry Gaussian Splatting for Feed-Forward Surround-view Driving Reconstruction

Junhong Lin, Kangli Wang, Shunzhou Wang, Songlin Fan, Ge Li, Wei Gao

机构 * Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology(广东省超高清沉浸媒体技术重点实验室) School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院) Peng Cheng Laboratory(鹏城实验室)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.CV

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18911 2025-10-23 physics.chem-ph cs.AI 57%

Prospects for Using Artificial Intelligence to Understand Intrinsic Kinetics of Heterogeneous Catalytic Reactions

Andrew J. Medford, Todd N. Whittaker, Bjarne Kreitz, David W. Flaherty, John R. Kitchin

机构 * organization= School of Chemical \& Biomolecular Engineering, Georgia Institute of Technology , addressline= 311 Ferst Drive NW , city= Atlanta , postcode= 30332 , state= GA , country= USA organization= Department of Chemical Engineering, Carnegie Mellon University , addressline= 5000 Forbes Street , city= Pittsburgh , postcode= 15213 , state= PA , country= USA

专题命中 端到端驾驶 :self-driving(abstract);分类 cs.AI

Comments Submitted to "Current Opinion in Chemical Engineering" for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12190 2025-10-15 cs.CV 57%

Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos

Shingo Yokoi, Kento Sasaki, Yu Yamaguchi

机构 * Turing Inc.(图灵公司)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.CV

Comments 2nd Place Winner, ICCV 2025 2COOOL Competition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11535 2025-10-14 cs.NI cs.AI 57%

A Flexible Multi-Agent Deep Reinforcement Learning Framework for Dynamic Routing and Scheduling of Latency-Critical Services

Vincenzo Norman Vitale, Antonia Maria Tulino, Andreas F. Molisch, Jaime Llorca

机构 * University of Naples Federico II(那不勒斯费德里科二世大学) University of Southern California(南加州大学) University of Trento(特伦托大学)

专题命中 端到端驾驶 :self-driving(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08046 2025-10-10 cs.AI 57%

LinguaSim: Interactive Multi-Vehicle Testing Scenario Generation via Natural Language Instruction Based on Large Language Models

Qingyuan Shi, Qingwen Meng, Hao Cheng, Qing Xu, Jianqiang Wang

机构 * National Key R&D Program of China(中华人民共和国国家重点研发计划)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13956 2025-09-18 cs.RO 57%

SEG-Parking: Towards Safe, Efficient, and Generalizable Autonomous Parking via End-to-End Offline Reinforcement Learning

Zewei Yang, Zengqi Peng, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Cheng Kar-Shun Robotics Institute, The Hong Kong University of Science and Technology(陈克逊机器人研究所,香港科技大学)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12437 2025-09-17 cs.AI 57%

Enhancing Physical Consistency in Lightweight World Models

Dingrui Wang, Zhexiao Sun, Zhouheng Li, Cheng Wang, Youlun Peng, Hongyuan Ye, Baha Zarrouki, Wei Li, Mattia Piccinini, Lei Xie, Johannes Betz

机构 * Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich(自动驾驶车辆系统教授职位,技术大学慕尼黑工程与设计学院) College of Control Science and Engineering, Zhejiang University(控制科学与工程学院,浙江大学) Nanjing University(南京大学) School of Mechatronics Engineering, Harbin Institute of Technology(机械电子工程学院,哈尔滨工业大学)

专题命中 端到端驾驶 :BEV(abstract);分类 cs.AI

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12068 2025-09-16 cs.CV 57%

End-to-End Learning of Multi-Organ Implicit Surfaces from 3D Medical Imaging Data

Farahdiba Zarin, Nicolas Padoy, Jérémy Dana, Vinkle Srivastav

机构 * University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France Institute of Image-Guided Surgery, IHU Strasbourg, Strasbourg, France Department of Diagnostic Radiology, McGill University, Montreal, Canada Augmented Intelligence \& Precision Health Laboratory (AIPHL), McGill University Health Centre Research Institute, Montreal, Canada

专题命中 端到端驾驶 :occupancy(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03238 2025-09-03 cs.RO 57%

RobotxR1: Enabling Embodied Robotic Intelligence on Large Language Models through Closed-Loop Reinforcement Learning

Liam Boyle, Nicolas Baumann, Paviththiren Sivasothilingam, Michele Magno, Luca Benini

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07611 2025-08-12 cs.RO 57%

End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy

Zifan Wang, Xun Yang, Jianzhuang Zhao, Jiaming Zhou, Teli Ma, Ziyao Gao, Arash Ajoudani, Junwei Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Human-Robot Interfaces and Interaction Lab., Istituto Italiano di Tecnologia, Italy(人机交互实验室,意大利理工学院)

专题命中 端到端驾驶 :LiDAR(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01506 2025-08-05 cs.LG cs.AI cs.PF 57%

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models

Zishan Shao, Yixiao Wang, Qinsi Wang, Ting Jiang, Zhixu Du, Hancheng Ye, Danyang Zhuo, Yiran Chen, Hai Li

机构 * Department of Statistical Science(统计科学系) Department of Electrical & Computer Engineering(电气与计算机工程系) Department of Computer Science(计算机科学系)

专题命中 端到端驾驶 :occupancy(abstract);分类 cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21371 2025-07-30 cs.CV 57%

Top2Pano: Learning to Generate Indoor Panoramas from Top-Down View

Zitong Zhang, Suranjan Gautam, Rui Yu

机构 * University of Louisville(路易斯维尔大学)

专题命中 端到端驾驶 :occupancy(abstract);分类 cs.CV

Comments ICCV 2025. Project page: https://top2pano.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17118 2025-07-24 cs.AI 57%

HySafe-AI: Hybrid Safety Architectural Analysis Framework for AI Systems: A Case Study

Mandar Pitale, Jelena Frtunikj, Abhinaw Priyadershi, Vasu Singh, Maria Spence

机构 * Nvidia Corporation(英伟达公司) Nvidia GmbH(英伟达德国公司)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.AI

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01729 2025-07-21 cs.CV 57%

PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth

Bu Jin, Weize Li, Baihan Yang, Zhenxin Zhu, Junpeng Jiang, Huan-ang Gao, Haiyang Sun, Kun Zhan, Hengtong Hu, Xueyang Zhang, Peng Jia, Hao Zhao

机构 * The Hong Kong University of Science and Technology(香港科技大学) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) Li Auto BAAI(北京人工智能研究院)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.CV

Comments Accepted at IEEE/RSJ IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23765 2025-07-18 cs.CV 57%

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Yun Li, Yiming Zhang, Tao Lin, Xiangrui Liu, Wenxiao Cai, Zheng Liu, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) China University of Geosciences(中国地质大学) Nanyang Technological University(南洋理工大学) BAAI(百度人工智能研究院) Stanford University(斯坦福大学)

专题命中 端到端驾驶 :autonomous driving(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01717 2025-07-09 cs.CV 57%

Driving View Synthesis on Free-form Trajectories with Generative Prior

Zeyu Yang, Zijie Pan, Yuankun Yang, Xiatian Zhu, Li Zhang

机构 * Fudan University(复旦大学) University of Surrey(Surrey大学)

专题命中 端到端驾驶 :end-to-end driving(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏