arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

共收录 7719 信号源:cs.CV, cs.GR, cs.RO

1. 点云 7719 篇

2603.04848 2026-03-13 cs.RO 57%

Hyperbolic Multiview Pretraining for Robotic Manipulation

双曲多视角预训练用于机器人操作

Jin Yang, Ping Wei, Yixin Chen, Nanning Zheng

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本文提出HyperMVP,一种基于双曲空间的多视角自监督预训练框架,通过GeoLink编码器学习结构化嵌入,提升机器人操作任务的性能。

Comments This paper was submitted to CVPR 2026 and was recommended for Findings, but the authors have withdrawn it and are currently adding more content to submit it elsewhere

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10871 2026-03-12 cs.RO 57%

FG-CLTP: Fine-Grained Contrastive Language Tactile Pretraining for Robotic Manipulation

FG-CLTP: 用于机器人操作的细粒度对比语言触觉预训练

Wenxuan Ma, Chaofan Zhang, Yinghao Cai, Guocai Yao, Shaowei Cui, Shuo Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 FG-CLTP通过细粒度对比学习提升机器人触觉感知,实现高精度分类和控制,减少误差并增强跨传感器泛化能力。

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10565 2026-03-12 cs.RO 57%

TacLoc: Global Tactile Localization on Objects from a Registration Perspective

TacLoc: 从配准角度实现物体的全局触觉定位

Zirui Zhang, Boyang Zhang, Fumin Zhang, Huan Yin

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 TacLoc通过单次点云配准方法实现物体的高效触觉定位,无需预训练模型或渲染数据,提升了姿态估计的准确性和效率。

Comments 8 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02558 2026-03-12 hep-ex cs.CV cs.LG 57%

Particle Trajectory Representation Learning with Masked Point Modeling

基于掩码点建模的粒子轨迹表示学习

Sam Young, Yeon-jae Jwa, Kazuhiro Terao

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出PoLAr-MAE,通过掩码点建模实现LArTPC图像的自监督学习,以高效学习物理轨迹表示,并发布大规模数据集促进后续研究。

Comments Preprint. 28 pages, 18 figures. v3 includes new results on data efficiency and attention maps

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09415 2026-03-11 cs.RO cs.AI 57%

From Flow to One Step: Real-Time Multi-Modal Trajectory Policies via Implicit Maximum Likelihood Estimation-based Distribution Distillation

从流到一步:通过隐式最大似然估计基于的分布蒸馏实现实时多模轨迹策略

Ju Dong, Liding Zhang, Lei Zhang, Yu Fu, Kaixin Bai, Zoltan-Csaba Marton, Zhenshan Bing, Zhaopeng Chen, Alois Christian Knoll, Jianwei Zhang

机构 * TAMS (Technical Aspects of Multimodal Systems), Department of Informatics, University of Hamburg(汉堡大学信息学院TAMS(多模态系统技术方面)) Technical University of Munich(慕尼黑技术大学) Agile Robots SE(敏捷机器人公司)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 通过隐式最大似然估计基于的分布蒸馏,实现实时多模轨迹策略,提升机器人操作的高频闭环控制与鲁棒性。

Comments https://sites.google.com/view/flow2one, 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09320 2026-03-11 cs.CV cs.AI 57%

SpaceSense-Bench: A Large-Scale Multi-Modal Benchmark for Spacecraft Perception and Pose Estimation

SpaceSense-Bench: 一个大规模多模态基准用于航天器感知与姿态估计

Aodi Wu, Jianhong Zuo, Zeyuan Zhao, Xubo Luo, Ruisuo Wang, Xue Wan

机构 * University of Chinese Academy of Sciences(中国科学院大学) Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心) Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 SpaceSense-Bench通过大规模多模态数据集提升航天器感知与姿态估计的鲁棒性与准确性。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00124 2026-03-11 cs.CV cs.AI 57%

OrthoAI: A Neurosymbolic Framework for Evidence-Grounded Biomechanical Reasoning in Clear Aligner Orthodontics

OrthoAI:一个用于清晰矫治器正畸的神经符号框架,用于基于证据的生物力学推理

Edouard Lansiaux, Margaux Leman, Mehdi Ammi

机构 * STaR-AI, Emergency Department, Lille University Hospital(STaR-AI急诊部,利尔大学医院) Artificial Intelligence and Data Semantics Laboratory, Paris 8 University(人工智能与数据语义实验室,巴黎第八大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 OrthOAI通过神经符号框架实现基于证据的生物力学推理,提升清晰矫治器正畸的自动化决策支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06572 2026-03-10 cs.CV cs.LG 57%

SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation

SCOPE: 场景上下文化增量少量样本3D分割

Vishal Thengane, Zhaochong An, Tianjin Huang, Son Lam Phung, Abdesselam Bouzerdoum, Lu Yin, Na Zhao, Xiatian Zhu

机构 * University of Surrey, UK(英国萨里大学) University of Wollongong, Australia(澳大利亚沃拉彭大学) University of Copenhagen, Denmark(丹麦哥本哈根大学) University of Exeter, UK(英国埃克塞特大学) Singapore University of Technology and Design, Singapore(新加坡科技与设计大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 SCOPE通过背景引导原型丰富框架,在3D分割中实现增量少量样本学习,提升新类别和平均IoU性能,同时减少遗忘。

Comments Accepted at CVPR 2026 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05760 2026-03-10 cs.RO cs.HC 57%

Task-Oriented Robot-Human Handovers on Legged Manipulators

面向任务的腿式机械臂人机交接

Andreea Tulbure, Carmen Scheidemann, Elias Steiner, Marco Hutter

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本文提出AFT-Handover框架,通过结合大语言模型和纹理转移实现零样本、可泛化的面向任务的人机交接,提升交接成功率和泛化能力。

Comments Accepted to 21st ACM/IEEE International Conference on Human-Robot Interaction (HRI) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05698 2026-03-10 cs.CV 57%

Point-based Instance Completion with Scene Constraints

基于场景约束的点云实例补全

Wesley Khademi, Li Fuxin

机构 * Oregon State University(俄勒冈州立大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出了一种基于点云的实例补全模型,通过引入场景约束和交叉注意力机制,实现对场景中任意尺度和姿态的对象的稳健补全。

Comments Published in ICLR 2025. Project Page: https://wkhademi.github.io/point_based_instance_completion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13172 2026-03-10 cs.CV 57%

WHU-STree: A Multi-modal Benchmark Dataset for Street Tree Inventory

WHU-STree: 一个用于街道树盘点的多模态基准数据集

Ruifei Ding, Zhe Chen, Wen Fan, Chen Long, Huijuan Xiao, Yelu Zeng, Zhen Dong, Bisheng Yang

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 WHU-STree是一个多模态的街道树盘点数据集,包含丰富的标注和多任务支持,用于提升城市街道树管理的自动化和智能化水平。

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing, 2026, 233: 519-542

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02808 2026-03-10 cs.RO cs.AI cs.SY eess.SY 57%

Improving the Resilience of Quadrotors in Underground Environments by Combining Learning-based and Safety Controllers

通过结合学习型控制器和安全控制器提高地下环境中四旋翼的鲁棒性

Isaac Ronald Ward, Mark Paral, Kristopher Riordan, Mykel J. Kochenderfer

机构 * Stanford Intelligent Systems Laboratory, Department of Aeronautics and Astronautics, Stanford University(斯坦福大学航空航天系)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本研究通过结合学习型和安全控制器,提高四旋翼在地下环境中的鲁棒性,实现任务完成与碰撞避免的平衡。

Comments Accepted and awarded best paper at the 11th International Conference on Control, Decision and Information Technologies (CoDIT 2025 - https://codit2025.org/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02858 2026-03-10 cs.CV 57%

Empowering Microscopic Traffic Simulators with Realistic Perception using Surrogate Sensor Models

通过使用替代传感器模型使微观交通模拟器具备现实感知能力

Tianheng Zhu, Yiheng Feng

机构 * Lyles School of Civil and Construction Engineering, Purdue University(普渡大学莱尔斯土木与建设工程学院)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 MIDAR通过替代传感器模型使微观交通模拟器具备现实感知能力,提升ITS应用的仿真真实性与效率。

Comments 27 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06250 2026-03-09 cs.CV 57%

Hierarchical Collaborative Fusion for 3D Instance-aware Referring Expression Segmentation

层次化协作融合用于3D实例感知指代表达分割

Keshen Zhou, Runnan Chen, Mingming Gong, Tongliang Liu

机构 * The University of Sydney(悉尼大学) The University of Melbourne(墨尔本大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 HCF-RES通过层次化视觉语义分解和渐进多级融合,实现了3D实例感知指代表达分割的高精度与细粒度定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05355 2026-03-09 cs.RO 57%

OmniDP: Beyond-FOV Large-Workspace Humanoid Manipulation with Omnidirectional 3D Perception

OmniDP: 在超广角大工作空间中实现人形机器人的全方位3D感知

Pei Qu, Zheng Li, Yufei Jia, Ziyun Liu, Liang Zhu, Haoang Li, Jinni Zhou, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 OmniDP通过360度全景点云感知和时间感知注意力池化机制,实现人形机器人在大工作空间中的稳健操作,优于传统深度相机方法。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01034 2026-03-09 cs.CV cs.AI cs.LG 57%

Reparameterized Tensor Ring Functional Decomposition for Multi-Dimensional Data Recovery

重新参数化张量环功能分解用于多维数据恢复

Yangyang Xu, Junbo Ke, You-Wei Wen, Chao Wang

机构 * Key Laboratory of Computing and Stochastic Mathematics (Ministry of Education)(计算与随机数学重点实验室(教育部)) School of Mathematics and Statistics(数学与统计学学院) Department of Statistics and Data Science(统计与数据科学系)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出了一种重新参数化的张量环功能分解方法,通过结合可学习的潜在张量和固定基底,提升多维数据恢复的性能。

Comments 22 pages, 18 figures, 12 tables. Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05111 2026-03-06 cs.RO cs.AI 57%

SPIRIT: Perceptive Shared Autonomy for Robust Robotic Manipulation under Deep Learning Uncertainty

SPIRIT:基于深度学习不确定性的感知共享自主控制以实现稳健的机器人操作

Jongseok Lee, Ribin Balachandran, Harsimran Singh, Jianxiang Feng, Hrishik Mishra, Marco De Stefano, Rudolph Triebel, Alin Albu-Schaeffer, Konstantin Kondak

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 SPIRIT通过结合深度学习的不确定性估计与触觉遥控,实现安全稳健的机器人操作,提升性能与可靠性。

Comments 19 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05017 2026-03-06 cs.RO 57%

Direct Contact-Tolerant Motion Planning With Vision Language Models

直接接触容忍的运动规划与视觉语言模型

He Li, Jian Sun, Chengyang Li, Guoliang Li, Qiyu Ruan, Shuai Wang, Chengzhong Xu

机构 * State Key Laboratory of Internet of Things for Smart City (SKL-IOTSC), University of Macau(物联网智能城市国家重点实验室,澳门大学) Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Department of Electrical and Computer Engineering, The University of Hong Kong(香港大学电子与计算机工程系)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本文提出直接接触容忍运动规划方法,利用视觉语言模型实现接触感知导航,提升机器人在拥挤环境中的鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06068 2026-03-06 cs.RO cs.AI 57%

MachaGrasp: Morphology-Aware Cross-Embodiment Dexterous Hand Articulation Generation for Grasping

MachaGrasp:基于形态的跨躯体灵巧手关节生成方法用于抓取

Heng Zhang, Kevin Yuchen Ma, Mike Zheng Shou, Weisi Lin, Yan Wu

机构 * Robotics & Autonomous Systems Division, Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR-I 2 R)(机器人与自主系统 division,信息通信研究所,科技研究局) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) Show Lab, National University of Singapore(Show Lab,国立新加坡大学)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 MachaGrasp提出一种基于形态的跨躯体灵巧手抓取生成方法,通过形态嵌入和特征抓取集生成抓取动作,实现高成功率的抓取任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17896 2026-03-05 cs.CV cs.AI 57%

EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations

EgoWorld:利用丰富的外参照观测将外参照视角转换为自身参照视角

Junho Park, Andrew Sangwoo Ye, Taein Kwon

机构 * AI Lab, LG Electronics(LG电子人工智能实验室) KAIST(韩国科学技术院) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 EgoWorld通过重建丰富的外参照观测来实现外参照到自身参照视角的转换,展示了在增强现实、虚拟现实和机器人应用中的先进性能和鲁棒性。

Comments Accepted by ICLR 2026. Project Page: https://redorangeyellowy.github.io/EgoWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02896 2026-03-04 cs.CV 57%

3D-DRES: Detailed 3D Referring Expression Segmentation

3D-DRES: 详细3D指称表达分割

Qi Chen, Changli Wu, Jiayi Ji, Yiwei Ma, Liujuan Cao

专题命中 点云 :3D vision(abstract);分类 cs.CV

AI总结 3D-DRES通过引入DetailRefer数据集和DetailBase架构,提升3D视觉语言理解的细粒度分割能力。

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01074 2026-03-03 cs.CV 57%

Adaptive Augmentation-Aware Latent Learning for Robust LiDAR Semantic Segmentation

自适应增强感知潜在学习用于鲁棒激光雷达语义分割

Wangkai Li, Zhaoyang Li, Yuwen Pan, Rui Sun, Yujia Chen, Tianzhu Zhang

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 A3Point通过自适应增强感知潜在学习框架,有效缓解恶劣天气下的语义偏移问题,提升激光雷达语义分割的鲁棒性。

Comments Accepted by International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00615 2026-03-03 cs.RO 57%

TGM-VLA: Task-Guided Mixup for Sampling-Efficient and Robust Robotic Manipulation

TGM-VLA:基于任务的混合学习用于高效且鲁棒的机器人操作

Fanqi Pu, Lei Jiang, Wenming Yang

机构 * Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(清华大学深圳国际研究生院) The National and Local Co-Build Humanoid Robotics Innovation Center(国家级与地方共建人形机器人创新中心)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 TGM-VLA通过优化关键帧采样策略和引入颜色反转投影模块,提升机器人操作任务的效率和鲁棒性。

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00486 2026-03-03 cs.CV 57%

Random Wins All: Rethinking Grouping Strategies for Vision Tokens

随机胜出:重新思考视觉token的分组策略

Qihang Fan, Yuang Ai, Huaibo Huang, Ran He

机构 * MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China(自动化研究所,中国科学院,北京) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(人工智能学院,中国科学院大学,北京)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出随机分组策略,通过简单方法提升视觉token处理效率,实验显示其在多种任务中表现优异。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00338 2026-03-03 cs.RO 57%

Layered Safety: Enhancing Autonomous Collision Avoidance via Multistage CBF Safety Filters

分层安全:通过多阶段CBF安全过滤器增强自主避障

Erina Yamaguchi, Ryan M. Bena, Gilbert Bahati, Aaron D. Ames

机构 * Caltech(卡内基梅隆大学)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本文提出一种多阶段CBF安全过滤器,通过预测和实时安全过滤提升机器人动态避障的鲁棒性和安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15020 2026-03-03 cs.RO 57%

ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision

ISS政策:具有隐式场景监督的可扩展扩散策略

Wenlong Xia, Jinhao Zhang, Ce Zhang, Yaojia Wang, Huizhe Li, Youmin Gong, Jie Mei

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 ISS策略通过隐式场景监督模块提升机器人操作的性能和鲁棒性,实现高效、泛化能力强的3D视觉-运动扩散策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23926 2026-03-03 cs.CV 57%

Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation

点-MoE:通过混合专家进行大规模多数据集训练用于3D语义分割

Xuweiyi Chen, Wentao Zhou, Aruni RoyChowdhury, Zezhou Cheng

机构 * University of Virginia(弗吉尼亚大学) MathWorks

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 Point-MoE通过混合专家方法在无需数据集标签的情况下,实现大规模多数据集联合训练,提升3D语义分割性能。

Comments Project page: https://point-moe.cs.virginia.edu/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15663 2026-03-03 cs.CV 57%

MSSPlace: Multi-Sensor Place Recognition with Visual and Text Semantics

MSSPlace: 多传感器位置识别与视觉和文本语义

Alexander Melekhin, Dmitry Yudin, Ilia Petryashin, Vitaly Bezuglyj

机构 * Intelligent Transport Laboratory, Moscow Institute of Physics and Technology(智能交通实验室,莫斯科物理技术学院) Artificial Intelligence Research Institute (AIRI)(人工智能研究机构(AIRI))

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 MSSPlace通过整合多传感器数据和视觉文本语义,提升位置识别性能,达到最先进的效果。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17297 2026-03-03 cs.CV cs.AI 57%

Towards Camera Open-set 3D Object Detection for Autonomous Driving Scenarios

面向自动驾驶场景的相机开放集3D目标检测

Zhuolin He, Xinrun Li, Jiacheng Tang, Shoumeng Qiu, Wenfu Wang, Xiangyang Xue, Jian Pu

机构 * ZILab

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 OS-Det3D通过两阶段框架提升自动驾驶中相机3D目标检测器对未知对象的发现与识别能力,同时提升已知对象的检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23945 2026-03-02 cs.CV cs.AI cs.MM 57%

PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning

PointCoT: 一种用于显式3D几何推理的多模态基准

Dongxu Zhang, Yiding Sun, Pengcheng Li, Yumou Liu, Hongqiang Lin, Haoran Xu, Xiaoxuan Mu, Liang Lin, Wenbiao Yan, Ning Yang, Chaowei Fang, Juanjuan Zhao, Jihua Zhu, Conghui He, Cheng Tan

机构 * Xi'an Jiaotong University(西安交通大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Institute of Automation, CASIA(中国科学院自动化研究所) Taiyuan University of Technology(太原理工大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 PointCoT通过显式链式推理提升3D几何推理能力,提出多模态基准和双流架构,实现对3D点云的高精度理解与推理。

详情

展开后加载摘要…

URL PDF HTML 收藏