arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

2026-02-11 至 2026-02-11 共收录 13 信号源:cs.CV, cs.GR, cs.RO

1. 三维重建 3 篇

2602.09918 2026-02-11 cs.CV cs.AI 79%

SARS: A Novel Face and Body Shape and Appearance Aware 3D Reconstruction System extends Morphable Models

SARS: 一种新型的面向面部和身体形状及外观的3D重建系统扩展可变形模型

Gulraiz Khan, Kenneth Y. Wertheim, Kevin Pimbblet, Waqas Ahmed

机构 * Centre of Excellence for Data Science, Artificial Intelligence and Modelling (DAIM), University of Hull, UK(数据科学、人工智能与建模卓越中心(DAIM)、赫尔大学)

专题命中 三维重建 :3D reconstruction(title,abstract);分类 cs.CV

AI总结 SARS提出了一种新型的3D重建系统,通过单张图像提取身体和面部信息,实现对全身体3D模型的精确重建,扩展了可变形模型的应用范围。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09278 2026-02-11 cs.CV cs.LG cs.RO 62%

UFM: A Simple Path towards Unified Dense Correspondence with Flow

UFM: 一种通往统一密集对应关系的简单路径:流

Yuchen Zhang, Nikhil Keetha, Chenwei Lyu, Bhuvan Jhamb, Yutian Chen, Yuheng Qiu, Jay Karhade, Shreyas Jha, Yaoyu Hu, Deva Ramanan, Sebastian Scherer, Wenshan Wang

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 三维重建 :3D reconstruction(abstract);分类 cs.CV、cs.RO

AI总结 UFM通过统一训练方法在密集对应关系任务中实现了更高的准确性和效率,优于专门方法。

Comments Project Page: https://uniflowmatch.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03200 2026-02-11 cs.CV cs.RO 62%

Transformer-Based Spatio-Temporal Association of Apple Fruitlets

基于变换器的苹果果粒时空关联

Harry Freeman, George Kantor

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 三维重建 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 本文提出基于变换器的苹果果粒时空关联方法,通过自注意力和交叉注意力机制提升小尺寸水果的关联精度,实验结果显示在商业果园中达到92.4%的F1分数。

Journal ref 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hangzhou, China, 2025, pp. 3018-3025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. NeRF 1 篇

2312.12122 2026-02-11 cs.CV cs.GR 62%

ZS-SRT: An Efficient Zero-Shot Super-Resolution Training Method for Neural Radiance Fields

ZS-SRT: 一种高效的零样本超分辨率训练方法用于神经辐射场

Xiang Feng, Yongbo He, Yubo Wang, Chengkai Wang, Zhenzhong Kuang, Jiajun Ding, Feiwei Qin, Jun Yu, Jianping Fan

机构 * School of Computer Science and Technology, Hangzhou Dianzi University(杭州电子科技大学计算机科学与技术学院) AI Lab at Lenovo Research(联想研究院人工智能实验室)

专题命中 NeRF :NeRF(abstract);分类 cs.CV、cs.GR

AI总结 ZS-SRT提出了一种无需高分辨率数据的NeRF超分辨率训练方法,通过内部学习和反向渲染提升训练效率与图像质量。

Journal ref Neurocomputing, Volume 590, 14 July 2024, Article 127714

详情

展开后加载摘要…

URL PDF HTML 收藏

3. Gaussian Splatting 4 篇

2602.09999 2026-02-11 cs.CV cs.GR 84%

Faster-GS: Analyzing and Improving Gaussian Splatting Optimization

Faster-GS:分析和改进高斯点云优化

Florian Hahlbohm, Linus Franke, Martin Eisemann, Marcus Magnor

机构 * Computer Graphics Lab, TU Braunschweig(图计算机图形实验室,图林根大学)

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);3DGS(abstract);分类 cs.CV、cs.GR

AI总结 Faster-GS通过整合和优化3DGS的核心策略,实现了训练速度提升五倍的同时保持视觉质量,为高斯点云优化提供了新的高效基准,并拓展至四维场景重建。

Comments Project page: https://fhahlbohm.github.io/faster-gaussian-splatting

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08625 2026-02-11 cs.CV 83%

OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics

OpenMonoGS-SLAM: 单目高斯点云SLAM与开放语义结合

Jisang Yoo, Gyeongjin Kang, Hyun-kyu Ko, Hyeonwoo Yu, Eunbyung Park

机构 * Sungkyunkwan University(成均馆大学) Yonsei University(延世大学)

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);3DGS(abstract);分类 cs.CV

AI总结 OpenMonoGS-SLAM通过结合3DGS与开放集语义,实现无需深度输入的单目SLAM,提升开放世界环境下的感知与建图性能。

Comments Work in progress. Project page: https://jisang1528.github.io/OpenMonoGS-SLAM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09816 2026-02-11 cs.CV 79%

CompSplat: Compression-aware 3D Gaussian Splatting for Real-world Video

CompSplat: 为真实世界视频设计的压缩感知3D高斯点散布

Hojun Song, Heejung Choi, Aro Kim, Chae-yeong Song, Gahyeon Kim, Soo Ye Kim, Jaehyup Lee, Sang-hyo Park

机构 * Kyungpook National University(Kyungpook国立大学) Adobe Research(Adobe研究院)

专题命中 Gaussian Splatting :Gaussian Splatting(title);novel view synthesis(abstract);分类 cs.CV

AI总结 CompSplat是一种针对真实世界视频压缩感知的3D高斯点散布框架,通过建模帧级压缩特性来减少帧间不一致性和累积几何误差,实现了在高压缩条件下更高质量的视角合成。

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09736 2026-02-11 cs.CV 70%

Toward Fine-Grained Facial Control in 3D Talking Head Generation

迈向3D说话人脸生成中的细粒度面部控制

Shaoyang Xie, Xiaofeng Cong, Baosheng Yu, Zhipeng Gui, Jie Gui, Yuan Yan Tang, James Tin-Yau Kwok

机构 * School of Cyber Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Lee Kong Chian School of Medicine, Nanyang Technological University(南洋理工大学李科金医学院) School of Remote Sensing and Information Engineering, Wuhan University(武汉大学遥感与信息工程学院) Purple Mountain Laboratories, Nanjing(南京紫金山实验室) Engineering Research Center of Blockchain Application, Supervision And Management (Southeast University), Ministry of Education(教育部区块链应用、监督与管理工程研究中心(东南大学)) Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 Gaussian Splatting :Gaussian Splatting(abstract);3DGS(abstract);分类 cs.CV

AI总结 本文提出FG-3DGS框架,通过频率感知解耦策略实现细粒度面部控制,提升3D说话人脸生成的精度和唇同步效果。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 点云 3 篇

2602.09414 2026-02-11 eess.SY cs.RO cs.SY 79%

Finite-time Stable Pose Estimation on TSE(3) using Point Cloud and Velocity Sensors

基于点云和速度传感器的有限时间稳定姿态估计(TSE(3))

Nazanin S. Hashkavaei, Abhijit Dongare, Neon Srinivasu, Amit K. Sanyal

机构 * Department of Mechanical \& Aerospace Engineering, Syracuse University, Syracuse, NY 13244

专题命中 点云 :point cloud(title,abstract);分类 cs.RO

AI总结 本文提出一种基于点云和速度传感器的有限时间稳定姿态估计方法,通过几何力学框架实现鲁棒性,适用于自主车辆。

Comments 17 pages, 8 figures, submitted to Automatica

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11565 2026-02-11 cs.CV 79%

SNAP: Towards Segmenting Anything in Any Point Cloud

SNAP:在任意点云中实现任意分割

Aniket Gupta, Hanhui Wang, Charles Saunders, Aruni RoyChowdhury, Hanumant Singh, Huaizu Jiang

机构 * Northeastern University(东北大学) MathWorks

专题命中 点云 :point cloud(title,abstract);分类 cs.CV

AI总结 SNAP提出一种统一的交互式3D点云分割模型,支持点和文本提示,实现跨领域泛化和高质量分割。

Comments Project Page, https://neu-vi.github.io/SNAP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09425 2026-02-11 cs.CV cs.LG 57%

Bridging the Modality Gap in Roadside LiDAR: A Training-Free Vision-Language Model Framework for Vehicle Classification

弥合道路激光雷达的模态差距:一种无需训练的视觉-语言模型框架用于车辆分类

Yiqiao Li, Bo Shang, Jie Wei

机构 * Department of Civil Engineering, City College of New York(城市学院土木工程系) Department of Computer Science, City College of New York(城市学院计算机科学系)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出一种无需训练的视觉-语言模型框架,用于解决道路激光雷达中稀疏点云与密集图像之间的模态差距问题,实现细粒度车辆分类。

Comments 12 pages, 10 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 空间理解 1 篇

2602.09638 2026-02-11 cs.CV 57%

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

VideoAfford: 通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding

Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA, Yiming Zhong, Wenti Yin, Yuhao Liu, Zhiqing Cui, Jiahao Yuan, Lu Dai, Zhiyuan Ma, Hui Xiong

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

AI总结 VideoAfford通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding,结合动态交互先验和空间感知损失函数,提升机器人操作的可操作区域识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 其他3D视觉 1 篇

2602.10115 2026-02-11 cs.CV 61%

Quantum Multiple Rotation Averaging

量子多旋转平均

Shuteng Wang, Natacha Kuete Meli, Michael Möller, Vladislav Golyanik

机构 * Max Planck Institute for Informatics, SIC(马克斯·普朗克信息研究所,SIC) University of Siegen(施普伦次大学)

专题命中 其他3D视觉 :3D vision(abstract,journal_ref);分类 cs.CV

AI总结 本文提出IQARS算法,通过量子退火器解决多旋转平均问题,实现更高的旋转同步精度。

Comments 16 pages, 13 figures, 4 tables; project page: https://4dqv.mpi-inf.mpg.de/QMRA/

Journal ref International Conference on 3D Vision (3DV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏