arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

2025-12-24 至 2025-12-24 共收录 12 信号源:cs.CV, cs.GR, cs.RO

1. Gaussian Splatting 2 篇

2512.20148 2025-12-24 cs.CV cs.RO 88%

Enhancing annotations for 5D apple pose estimation through 3D Gaussian Splatting (3DGS)

通过3D高斯点绘(3DGS)增强5D苹果姿态估计的标注

Robert van de Ven, Trim Bresilla, Bram Nelissen, Ard Nieuwenhuizen, Eldert J. van Henten, Gert Kootstra

机构 * organization= Agricultural Biosystems Engineering, Wageningen University \& Research , addressline= Droevendaalsesteeg 1 , city= Wageningen , postcode= 6708 PB , country= the Netherlands organization= Agrosystems Research, Wageningen University \& Research , addressline= Droevendaalsesteeg 1 , city= Wageningen , postcode= 6708 PB , country= the Netherlands

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);3DGS(title);3D reconstruction(abstract);分类 cs.CV、cs.RO

AI总结 通过3D高斯点绘技术,简化苹果姿态估计的标注流程,提升训练数据效率与模型性能。

Comments 33 pages, excluding appendices. 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20495 2025-12-24 cs.AR 82%

Nebula: Enable City-Scale 3D Gaussian Splatting in Virtual Reality via Collaborative Rendering and Accelerated Stereo Rasterization

Nebula: 通过协作渲染和加速立体光栅化实现城市级3D高斯散射

He Zhu, Zheng Liu, Xingyang Li, Anbang Wu, Jieru Zhao, Fangxin Liu, Yiming Gan, Jingwen Leng, Yu Feng

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);3DGS(abstract)

AI总结 Nebula通过协作渲染和加速立体光栅化,实现城市级3D高斯散射的高效渲染,显著降低带宽需求并提升VR体验。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 点云 5 篇

2503.07940 2025-12-24 cs.CV cs.RO eess.IV 81%

BUFFER-X: Towards Zero-Shot Point Cloud Registration in Diverse Scenes

BUFFER-X:迈向多样化场景下的零样本点云配准

Minkyun Seo, Hyungtae Lim, Kanghee Lee, Luca Carlone, Jaesik Park

机构 * Laboratory for Information & Decision Systems(信息与决策系统实验室) Massachusetts Institute of Technology(麻省理工学院)

专题命中 点云 :point cloud(title,abstract);分类 cs.CV、cs.RO

AI总结 BUFFER-X通过自适应体素大小、最远点采样和补丁尺度归一化,实现多样化场景下的零样本点云配准,无需先验信息或手动调参。

Comments 20 pages, 14 figures. Accepted as a highlight paper at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20217 2025-12-24 cs.CV 57%

LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation

LiteFusion: 从基于视觉到多模态的3D目标检测器的最小适应

Xiangxuan Ren, Zhongdao Wang, Pin Tang, Guoqing Wang, Jilai Zheng, Chao Ma

机构 * China Ministry of Education (MOE) Key Laboratory of Artificial Intelligence(中国教育部人工智能重点实验室) Artificial Intelligence Institute, Shanghai Jiao Tong University(上海交通大学人工智能学院) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 LiteFusion通过将LiDAR数据作为补充几何信息源,无需专用LiDAR编码器,显著提升多模态3D目标检测性能。

Comments 13 pages, 9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17085 2025-12-24 cs.RO cs.LG 57%

Deformable Cluster Manipulation via Whole-Arm Policy Learning

通过全身政策学习实现可变形簇 manipulation

Jayadeep Jacob, Wenzheng Zhang, Houston Warren, Paulo Borges, Tirthankar Bandyopadhyay, Fabio Ramos

机构 * School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) Data61, CSIRO(CSIRO数据61研究所) Orica(奥里卡公司) NVIDIA Corporation(NVIDIA公司)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本文提出了一种基于全身政策学习的方法,通过整合3D点云和本体感觉触觉信息,实现对可变形簇的高效操纵,并在输电线路清除任务中展示了零样本迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05634 2025-12-24 cs.CV cs.LG eess.IV 57%

Fusarium head blight detection, spikelet estimation, and severity assessment in wheat using 3D convolutional neural networks

利用3D卷积神经网络检测小麦 Fusarium head blight、估计穗数及评估病情严重程度

Oumaima Hamila, Christopher J. Henry, Oscar I. Molina, Christopher P. Bidinosti, Maria Antonia Henriquez

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本研究利用3D卷积神经网络对小麦FHB进行检测、穗数估计及严重程度评估,实现100%检测准确率和高精度估计

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19743 2025-12-24 cs.LG cs.AI cs.CG 50%

From Theory to Throughput: CUDA-Optimized APML for Large-Batch 3D Learning

从理论到吞吐量:用于大规模批量3D学习的CUDA优化APML

Sasan Sharifipour, Constantino Álvarez Casado, Manuel Lage Cañellas, Miguel Bordallo López

机构 * Center for Machine Vision and Signal Analysis (CMVS), University of Oulu(机器视觉与信号分析中心(CMVS),奥卢大学)

专题命中 点云 :point cloud(abstract)

AI总结 CUDA-APML通过稀疏GPU实现优化大规模批量3D学习的APML,显著降低内存使用并保持模型精度。

Comments 5 pages, 2 figures, 2 tables, 5 formulas, 34 references, journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 新视角合成 1 篇

2512.20107 2025-12-24 cs.CV 57%

UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View Synthesis

UMAMI:统一掩码自回归模型和确定性渲染用于视图合成

Thanh-Tung Le, Tuan Pham, Tung Nguyen, Deying Kong, Xiaohui Xie, Stephan Mandt

机构 * UCI(加州大学伯克利分校) UCLA(加州大学洛杉矶分校) Google(谷歌)

专题命中 新视角合成 :novel view synthesis(abstract);分类 cs.CV

AI总结 UMAMI通过结合掩码自回归模型和确定性渲染,实现了视图合成中图像质量与渲染效率的统一。

Comments Accepted to NeurIPS 2025. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 空间理解 3 篇

2512.20557 2025-12-24 cs.CV 79%

Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models

在4D中学习推理:面向视觉语言模型的动态空间理解

Shengchao Zhou, Yuxin Chen, Yuying Ge, Wei Huang, Jiehong Lin, Ying Shan, Xiaojuan Qi

机构 * The University of Hong Kong(香港大学) ARC Lab, Tencent PCG(腾讯PCG实验室)

专题命中 空间理解 :spatial understanding(title);point cloud(abstract);分类 cs.CV

AI总结 本研究提出DSR套件,通过生成多选题-答案对和轻量级几何选择模块提升视觉语言模型的动态空间推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20501 2025-12-24 cs.CV 57%

Bridging Modalities and Transferring Knowledge: Enhanced Multimodal Understanding and Recognition

弥合模态与知识转移:增强多模态理解和识别

Gorjan Radevski

机构 * University of Bristol(布里斯托大学) University of Würzburg(乌尔姆大学) Processing Speech and Images (PSI)(语音与图像处理组)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 本文提出多模态对齐、翻译、融合和转移方法,提升复杂输入的理解与识别能力,涵盖空间语言、医学文本、知识图谱和动作识别等多个领域。

Comments Ph.D. manuscript; Supervisors/Mentors: Marie-Francine Moens and Tinne Tuytelaars

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13387 2025-12-24 cs.CV eess.IV 57%

From Binary to Semantic: Utilizing Large-Scale Binary Occupancy Data for 3D Semantic Occupancy Prediction

从二元到语义:利用大规模二元占用数据进行3D语义占用预测

Chihiro Noguchi, Takaki Yamamoto

机构 * InfoTech, Toyota Motor Corporation(丰田汽车公司信息科技部)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

AI总结 本文提出一种基于二元占用数据的框架,通过预训练和自动标注方法提升3D语义占用预测的性能。

Comments Accepted to ICCV Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. SLAM与定位 1 篇

2507.20538 2025-12-24 cs.RO 57%

Uni-Mapper: Unified Mapping Framework for Multi-modal LiDARs in Complex and Dynamic Environments

Uni-Mapper:多模态激光雷达在复杂动态环境中的统一映射框架

Gilhwan Kang, Hogyun Kim, Byunghee Choi, Seokhwan Jeong, Young-Sik Shin, Younggun Cho

机构 * Hyundai Motor Company(现代汽车公司) Inha University(inha大学) Korea Institute of Machinery and Materials(韩国机械材料研究院)

专题命中 SLAM与定位 :point cloud(abstract);分类 cs.RO

AI总结 Uni-Mapper通过动态感知和多模态激光雷达融合技术,实现复杂动态环境下的统一地图构建与回环检测。

Comments 18 pages, 14 figures

Journal ref 2025

详情

展开后加载摘要…

URL PDF HTML 收藏