arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

自动驾驶

自动驾驶感知、规划、BEV、占用预测、激光雷达和仿真评测。

共收录 6072 信号源:cs.RO, cs.CV, eess.IV, cs.AI

1. 感知 6072 篇

2509.26324 2026-03-03 cs.RO cs.AI cs.MA 62%

COMRES-VLM: Coordinated Multi-Robot Exploration and Search using Vision Language Models

COMRES-VLM: 基于视觉语言模型的多机器人协同探索与搜索

Ruiyang Wang, Hao-Lun Hsu, David Hunt, Jiwoo Kim, Shaocheng Luo, Miroslav Pajic

专题命中 感知 :occupancy(abstract);分类 cs.RO、cs.AI

AI总结 COMRES-VLM通过视觉语言模型实现多机器人系统的协同探索与搜索,提升探索效率和目标物体搜索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22731 2026-02-27 cs.RO cs.CV 62%

Sapling-NeRF: Geo-Localised Sapling Reconstruction in Forests for Ecological Monitoring

Sapling-NeRF: 用于森林生态监测的地理定位幼树重建

Miguel Ángel Muñoz-Bañón, Nived Chebrolu, Sruthi M. Krishna Moorthy, Yifu Tao, Fernando Torres, Roberto Salguero-Gómez, Maurice Fallon

机构 * Oxford Robotics Institute, Department of Engineering Science, University of Oxford, Oxford, UK(牛津大学机器人研究所、工程科学系、牛津大学、牛津、英国) Group of Automation, Robotics and Computer Vision, University of Alicante, Alicante, Spain(自动化、机器人与计算机视觉小组、阿尔瓦登特大学、阿尔瓦登特、西班牙) Department of Biology, University of Oxford, Oxford, UK(生物学系、牛津大学、牛津、英国)

专题命中 感知 :LiDAR(abstract);分类 cs.RO、cs.CV

AI总结 本文提出Sapling-NeRF方法,结合NeRF、LiDAR SLAM和GNSS,实现地理定位的幼树重建,提升森林生态监测的精度和长期数据采集能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06491 2026-02-27 cs.CV cs.RO 62%

PPT: Pretraining with Pseudo-Labeled Trajectories for Motion Forecasting

PPT:基于伪标签轨迹的运动预测预训练

Yihong Xu, Yuan Yin, Éloi Zablocki, Tuan-Hung Vu, Alexandre Boulch, Matthieu Cord

专题命中 感知 :autonomous driving(abstract);分类 cs.RO、cs.CV

AI总结 PPT通过利用伪标签轨迹进行预训练,提升运动预测的泛化能力和在低数据环境下的性能。

Comments 8 pages, 6 figures, accepted to ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18207 2026-02-27 cs.CV cs.AI 62%

From Open Vocabulary to Open World: Teaching Vision Language Models to Detect Novel Objects

从开放词汇到开放世界:教会视觉语言模型检测新物体

Zizhao Li, Zhengkang Xiang, Joseph West, Kourosh Khoshelham

机构 * The University of Melbourne Parkville, VIC, Australia(墨尔本大学帕克维尔分校)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种开放世界框架,使OVD模型能够检测新物体,通过引入OWEL和MSCAL方法提升模型对远超出分布物体的识别能力。

Comments Accepted by BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18850 2026-02-24 cs.RO cs.AI 62%

When the Inference Meets the Explicitness or Why Multimodality Can Make Us Forget About the Perfect Predictor

当推理遇见显式性或为什么多模态可以让我们忘记完美预测器

J. E. Domínguez-Vidal, Alberto Sanfeliu

机构 * Institut de Robòtica i Informàtica Industrial (CSIC-UPC)(机器人与信息工业研究所(CSIC-UPC)) Universitat Politècnica de Catalunya - BarcelonaTech (UPC)(加泰罗尼亚理工大学-巴塞罗那技术大学(UPC))

专题命中 感知 :LiDAR(abstract);分类 cs.RO、cs.AI

AI总结 本文探讨了在人机协作任务中,显式通信与推理系统结合的效果,发现人类更倾向于自然的交互方式,且最佳策略是两者的结合。

Comments Original version submitted to the International Journal of Social Robotics. Final version available on the SORO website

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00497 2026-02-24 cs.RO cs.CV 62%

FLUID: A Fine-Grained Lightweight Urban Signalized-Intersection Dataset of Dense Conflict Trajectories

FLUID: 一种细粒度轻量级密集冲突轨迹的城市信号交叉口数据集

Yiyang Chen, Zhigang Wu, Guohong Zheng, Xuesong Wu, Liwen Xu, Haoyuan Tang, Zhaocheng He, Haipeng Zeng

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(智能系统工程学院,中山大学) Guangdong Provincial Key Laboratory of Intelligent Transportation Systems(广东省智能交通系统重点实验室) Pengcheng Laboratory(鹏城实验室)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO、cs.CV

AI总结 FLUID数据集通过细粒度轨迹记录和轻量级处理框架,为研究城市信号交叉口的密集冲突提供了高时空精度的数据支持。

Comments 30 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20174 2026-02-24 cs.CV cs.AI 62%

LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks

LRR-Bench:左、右还是旋转?视觉-语言模型在空间理解任务上仍面临挑战

Fei Kong, Jinhao Duan, Kaidi Xu, Zhenhua Guo, Xiaofeng Zhu, Xiaoshuang Shi

机构 * University of Electronic Science and Technology of China(电子科技大学) Drexel University(德雷塞尔大学) Tianyijiaotong Technology Ltd.(天奕交通技术有限公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 LRR-Bench通过合成数据评估视觉-语言模型在空间理解任务上的性能,发现其在复杂任务上的表现远低于人类水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17252 2026-02-20 cs.CV cs.SY eess.IV eess.SY 62%

A Multi-modal Detection System for Infrastructure-based Freight Signal Priority

基于基础设施的货运信号优先的多模态检测系统

Ziyan Zhang, Chuheng Wei, Xuanpeng Zhao, Siyan Li, Will Snyder, Mike Stas, Peng Hao, Kanok Boriboonsomsin, Guoyuan Wu

专题命中 感知 :LiDAR(abstract);分类 cs.CV、eess.IV

AI总结 本文提出了一种基于激光雷达和摄像头的多模态货运车辆检测系统,通过混合传感架构和卡尔曼滤波实现稳定实时性能,用于支持基于基础设施的货运信号优先应用。

Comments 12 pages, 15 figures. Accepted at ICTD 2026. Final version to appear in ASCE Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01423 2026-02-17 cs.CV cs.LG cs.RO 62%

3DRot: Rediscovering the Missing Primitive for RGB-Based 3D Augmentation

3DRot:重新发现RGB基于3D增强中缺失的基本原理

Shitian Yang, Deyu Li, Xiaoke Jiang, Lei Zhang

机构 * International Digital Economy Academy (IDEA), Shenzhen, China(国际数字经济学院(IDEA)) Shenzhen University, Shenzhen, China(深圳大学)

专题命中 感知 :LiDAR(abstract);分类 cs.RO、cs.CV

AI总结 3DRot是一种新的RGB基于3D增强方法,通过旋转和镜像图像并更新相关参数,提升3D检测和深度估计的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12902 2026-02-16 cs.CV cs.AI cs.LG cs.SE 62%

Robustness of Object Detection of Autonomous Vehicles in Adverse Weather Conditions

自动驾驶车辆在恶劣天气条件下的目标检测鲁棒性

Fox Pettersen, Hong Zhu

机构 * School of Engineering, Computing and Mathematics, Oxford Brookes University(工程、计算与数学学院,奥克斯伯勒斯大学)

专题命中 感知 :self-driving(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种评估自动驾驶车辆在恶劣天气条件下目标检测鲁棒性的方法,通过数据增强生成合成数据,测试不同模型在恶劣条件下的表现,发现Faster R-CNN模型鲁棒性最高。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10943 2026-02-12 cs.CV cs.RO 62%

Towards Learning a Generalizable 3D Scene Representation from 2D Observations

从2D观测向量学习通用的3D场景表示

Martin Gromniak, Jan-Gerrit Habekost, Sebastian Kamp, Sven Magg, Stefan Wermter

机构 * University of Hamburg - Department of Informatics(汉堡大学信息学院) ZAL Center of Applied Aeronautical Research(应用航空研究中心) Hamburger Informatik Technologie-Center e.V. (HITeC)(汉堡信息科技中心(HITeC))

专题命中 感知 :occupancy(abstract);分类 cs.RO、cs.CV

AI总结 本文提出了一种通用神经辐射场方法,通过第一人称机器人观测学习3D场景表示,实现了对未见场景的泛化能力,并在真实场景中验证了其高精度的3D重建性能。

Comments Paper accepted at ESANN 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21164 2026-02-10 cs.CV cs.AI cs.LG 62%

Adversarial Wear and Tear: Exploiting Natural Damage for Generating Physical-World Adversarial Examples

对抗性磨损:利用自然损坏生成物理世界对抗性示例

Samra Irshad, Seungkyu Lee, Nassir Navab, Hong Joo Lee, Seong Tae Kim

机构 * Department of Computer Science and Engineering, Kyung Hee University(计算机科学与工程系,庆熙大学) Technical University of Munich(慕尼黑技术大学) Seoul National University of Science and Technology(首尔科学理工大学) G-LAMP NEXUS Institute, Kyung Hee University(G-LAMP NEXUS研究所,庆熙大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 本文提出AdvWT,通过模拟自然磨损生成更真实有效的物理世界对抗性示例,提升深度神经网络在现实环境中的鲁棒性。

Comments Accepted to IEEE Transactions in Secure and Dependable Computing. This version corresponds to the author's accepted manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07051 2026-02-10 cs.CV cs.AI cs.LG cs.NE 62%

Neural Sentinel: Unified Vision Language Model (VLM) for License Plate Recognition with Human-in-the-Loop Continual Learning

神经哨兵:用于车牌识别的统一视觉语言模型(VLM)与人类在循环持续学习

Karthik Sivakoti

专题命中 感知 :occupancy(abstract);分类 cs.CV、cs.AI

AI总结 Neural Sentinel利用视觉语言模型实现车牌识别与多任务学习,提升准确率并减少系统复杂性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04692 2026-02-09 cs.CV cs.AI 62%

DRMOT: A Dataset and Framework for RGBD Referring Multi-Object Tracking

DRMOT:用于RGBD参照多目标跟踪的数据集和框架

Sijia Chen, Lijuan Ma, Yanqiu Yu, En Yu, Liman Liu, Wenbing Tao

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 本文提出DRMOT任务,通过融合RGB、深度和语言信息,构建DRSet数据集并提出DRTrack框架,提升多目标跟踪的3D感知能力。

Comments https://github.com/chen-si-jia/DRMOT

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05966 2026-02-06 cs.CV cs.AI 62%

LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation

LSA:局部语义对齐用于增强交通视频生成中的时间一致性

Mirlan Karimov, Teodora Spasojevic, Markus Braun, Julian Wiederer, Vasileios Belagiannis, Marc Pollefeys

机构 * Mercedes-Benz AG(梅赛德斯-奔驰集团) ETH Zurich(苏黎世联邦理工学院) Friedrich-Alexander University Erlangen-Nuremberg(埃朗根-纽伦堡弗里德里希-亚历山大大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 LSA通过局部语义对齐提升交通视频生成的时间一致性,无需外部控制信号。

Comments Accepted to IEEE IV 2026. 8 pages, 3 figures. Code available at https://github.com/mirlanium/LSA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05557 2026-02-06 cs.CV cs.RO 62%

PIRATR: Parametric Object Inference for Robotic Applications with Transformers in 3D Point Clouds

PIRATR:基于变换器的参数化物体推理用于机器人应用的3D点云

Michael Schwingshackl, Fabio F. Oberweger, Mario Niedermeyer, Huemer Johannes, Markus Murschitz

机构 * AIT Austrian Institute of Technology Center for Vision, Automation & Control(奥地利技术研究院视觉、自动化与控制中心)

专题命中 感知 :LiDAR(abstract);分类 cs.RO、cs.CV

AI总结 PIRATR通过端到端3D点云检测框架,结合变换器实现参数化物体的6自由度姿态和属性估计,适用于机器人应用。

Comments 8 Pages, 11 Figures, Accepted at 2026 IEEE International Conference on Robotics & Automation (ICRA) Vienna

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03064 2026-02-04 cs.CV cs.AI 62%

JRDB-Pose3D: A Multi-person 3D Human Pose and Shape Estimation Dataset for Robotics

JRDB-Pose3D: 一个多人群3D人体姿态和形状估计数据集用于机器人

Sandika Biswas, Kian Izadpanah, Hamid Rezatofighi

机构 * Monash University(莫纳什大学) Sharif University(沙里夫大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 JRDB-Pose3D数据集通过移动机器人平台捕捉多人群的室内外环境,提供丰富的3D人体姿态和形状注释,用于机器人感知和人机交互等应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12968 2026-02-04 cs.CV cs.RO 62%

OptiPMB: Enhancing 3D Multi-Object Tracking with Optimized Poisson Multi-Bernoulli Filtering

OptiPMB:通过优化的泊松多伯努利滤波增强3D多目标跟踪

Guanhua Ding, Yuxuan Xia, Runwei Guan, Qinchen Wu, Tao Huang, Weiping Ding, Jinping Sun, Guoqiang Mao

机构 * School of Electronic Information Engineering, Beihang University(电子信息工程学院,北京航空航天大学) School of Automation and Intelligent Sensing, Shanghai Jiaotong University(自动化与智能感知学院,上海交通大学) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科技大学) College of Science and Engineering, James Cook University(科学与工程学院,詹姆斯库克大学) School of Artificial Intelligence and Computer Science, Nantong University(人工智能与计算机科学学院,南通大学) Faculty of Data Science, City University of Macau(数据科学学院,澳门城市大学) Research Laboratory of Smart Driving and Intelligent Transportation Systems, Southeast University(智能驾驶与智能交通系统研究实验室,东南大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO、cs.CV

AI总结 OptiPMB通过优化的泊松多伯努利滤波器和创新设计提升3D多目标跟踪的准确性与性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12638 2026-02-03 cs.CV cs.AI 62%

Mixed Precision PointPillars for Efficient 3D Object Detection with TensorRT

混合精度PointPillars用于基于TensorRT的高效3D目标检测

Ninnart Fuengfusin, Keisuke Yoneda, Naoki Suganuma

机构 * Advanced Mobility Research Institute(先进移动研究院) Kanazawa University(金泽大学)

专题命中 感知 :LiDAR(abstract);分类 cs.CV、cs.AI

AI总结 本文提出混合精度PointPillars框架,结合TensorRT实现高效3D目标检测,通过量化感知训练和贪心搜索优化模型性能,降低延迟达2.538倍。

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17895 2026-01-27 cs.CV cs.RO 62%

Masked Depth Modeling for Spatial Perception

遮蔽深度建模用于空间感知

Bin Tan, Changjiang Sun, Xiage Qin, Hanat Adai, Zelin Fu, Tianxiang Zhou, Han Zhang, Yinghao Xu, Xing Zhu, Yujun Shen, Nan Xue

专题命中 感知 :autonomous driving(abstract);分类 cs.RO、cs.CV

AI总结 LingBot-Depth通过遮蔽深度建模和自动化数据整理流水线,在深度精度和像素覆盖上优于顶级RGB-D相机,提供跨RGB和深度模态的对齐潜在表示。

Comments Tech report, 19 pages, 15 figures and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09952 2026-01-16 cs.CV cs.RO 62%

OT-Drive: Out-of-Distribution Off-Road Traversable Area Segmentation via Optimal Transport

OT-Drive: 基于最优传输的离群分布越野可通行区域分割

Zhihua Zhao, Guoqiang Li, Chen Min, Kangping Lu

机构 * School of Mechanical Engineering, Beijing Institute of Technology(北京理工大学机械工程学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Shandong Pengxiang Automobile Co., Ltd(山东鹏翔汽车有限公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO、cs.CV

AI总结 OT-Drive通过最优传输方法实现离群分布越野可通行区域分割,提升自动驾驶在复杂环境下的鲁棒性和泛化能力。

Comments 9 pages, 8 figures, 6 tables. This work has been submitted to the IEEE for possible publication. Code will be released upon acceptance

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07718 2026-01-13 cs.RO cs.AI 62%

Hiking in the Wild: A Scalable Perceptive Parkour Framework for Humanoids

野外徒步:一种可扩展的感知跳跃框架用于人形机器人

Shaoting Zhu, Ziwen Zhuang, Mengjie Zhao, Kun-Ying Lee, Hang Zhao

机构 * IIIS, Tsinghua University(清华大学人工智能研究院) Shanghai Qi Zhi Institute(上海启智研究院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)

专题命中 感知 :LiDAR(abstract);分类 cs.RO、cs.AI

AI总结 本文提出了一种可扩展的感知跳跃框架,通过结合地形边缘检测和脚部体积点,实现人形机器人在复杂地形中的稳健徒步。

Comments Project Page: https://project-instinct.github.io/hiking-in-the-wild

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04404 2026-01-09 cs.CV cs.AI 62%

3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation

3D-Agent:三模态多智能体协作用于可扩展的3D物体标注

Jusheng Zhang, Yijia Fan, Zimo Wen, Jian Wang, Keze Wang

机构 * Sun Yat-sen University(中山大学) Shanghai Jiao Tong University(上海交通大学) Snap Inc.(Snap公司)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 Tri MARF通过三模态多智能体协作提升大规模3D物体标注效率与精度

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01386 2026-01-06 cs.CV cs.AI 62%

ParkGaussian: Surround-view 3D Gaussian Splatting for Autonomous Parking

ParkGaussian:用于自动驾驶停车的环绕视图3D高斯点云法

Xiaobao Wei, Zhangjie Ye, Yuxiang Gu, Zunjie Zhu, Yunfei Guo, Yingying Shen, Shan Zhao, Ming Lu, Haiyang Sun, Bing Wang, Guang Chen, Rongfeng Lu, Hangjun Ye

机构 * Hangzhou Dianzi University(杭州电子科技大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 ParkGaussian通过整合3D高斯点云法,实现了停车场景的高质量重建,提升了自动驾驶停车任务中的感知一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24111 2026-01-01 cs.CV cs.RO 62%

Guided Diffusion-based Generation of Adversarial Objects for Real-World Monocular Depth Estimation Attacks

引导扩散生成对抗性物体用于现实世界单目深度估计攻击

Yongtao Chen, Yanbo Wang, Wentao Zhao, Guole Shen, Tianchen Deng, Jingchuan Wang

机构 * School of Automation and Intelligent Sensing, Institute of Medical Robotics, Shanghai Jiao Tong University(自动化与智能感知学院、医疗机器人研究院、上海交通大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO、cs.CV

AI总结 本文提出一种无需训练的生成对抗攻击框架,通过扩散模型生成自然且场景一致的对抗性物体,以提升自动驾驶系统中单目深度估计的对抗攻击效果和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17821 2026-01-01 cs.RO cs.AI cs.LG 62%

CAML: Collaborative Auxiliary Modality Learning for Multi-Agent Systems

CAML: 多智能体系统中的协同辅助模态学习

Rui Liu, Yu Shen, Peng Gao, Pratap Tokekar, Ming Lin

机构 * University of Maryland, College Park(马里兰大学 College Park 分校) Adobe Research(Adobe 研究院) North Carolina State University(北卡罗来纳州立大学)

专题命中 感知 :autonomous driving(abstract);分类 cs.RO、cs.AI

AI总结 CAML通过多智能体协作和共享多模态数据,提升多模态学习在资源受限环境下的性能,实现事故检测和语义分割的显著改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12843 2026-01-01 cs.CV cs.AI 62%

SLTNet: Efficient Event-based Semantic Segmentation with Spike-driven Lightweight Transformer-based Networks

SLTNet: 面向事件感知语义分割的高效轻量级脉冲神经网络

Xianlei Long, Xiaxin Zhu, Fangming Guo, Wanyi Zhang, Qingyi Gu, Chao Chen, Fuqiang Gu

机构 * College of Computer Science, Chongqing University(重庆大学计算机学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 SLTNet通过脉冲驱动轻量级变压器网络实现高效事件感知语义分割,提升性能并降低能耗

Comments Accepted by IROS 2025 (2025 IEEE/RSJ International Conference on Intelligent Robots and Systems)

Journal ref 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23208 2025-12-30 cs.CV cs.AI 62%

Exploring Syn-to-Real Domain Adaptation for Military Target Detection

探索合成到现实领域适应用于军事目标检测

Jongoh Jeong, Youngjin Oh, Gyeongrae Nam, Jeongeun Lee, Kuk-Jin Yoon

机构 * Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院机械工程系) Drone Research and Development Laboratory, LIG Nex1(LIG Nex1无人机研究与发展实验室)

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、cs.AI

AI总结 本文提出使用Unreal Engine生成RGB合成数据,以提升军事目标检测在跨领域适应中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21975 2025-12-29 eess.IV cs.CV 62%

RT-Focuser: A Real-Time Lightweight Model for Edge-side Image Deblurring

RT-Focuser: 一种用于边缘侧图像去模糊的轻量级模型

Zhuoyu Wu, Wenhui Ou, Qiawei Zheng, Jiayan Yang, Quanjun Wang, Wenqi Fang, Zheng Wang, Yongkui Yang, Heshan Li

专题命中 感知 :autonomous driving(abstract);分类 cs.CV、eess.IV

AI总结 RT-Focuser是一种轻量级实时图像去模糊模型,通过三个关键模块提升速度与精度,实现高PSNR和高帧率,适用于边缘计算场景。

Comments 2 pages, 2 figures, this paper already accepted by IEEE ICTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21641 2025-12-29 cs.CV cs.AI 62%

TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References

TrackTeller: 基于时间的多模态3D定位用于行为依赖的对象引用

Jiahong Yu, Ziqi Wang, Hailiang Zhao, Wei Zhai, Xueqiang Yan, Shuiguang Deng

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学) Huawei Technologies Ltd.(华为技术有限公司)

专题命中 感知 :LiDAR(abstract);分类 cs.CV、cs.AI

AI总结 TrackTeller通过融合LiDAR图像、语言条件解码和时间推理,提升动态3D场景中基于语言的物体定位与跟踪性能。

详情

展开后加载摘要…

URL PDF HTML 收藏