arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

2026-01-13 至 2026-01-13 共收录 30 信号源:cs.CV, cs.GR, cs.RO

1. 三维重建 5 篇

2601.07335 2026-01-13 cs.CV 57%

Reconstruction Guided Few-shot Network For Remote Sensing Image Classification

基于重建的少样本网络用于遥感图像分类

Mohit Jaiswal, Naman Jain, Shivani Pathak, Mainak Singha, Nikunja Bihari Kar, Ankit Jha, Biplab Banerjee

机构 * Dept. of CSE, The LNMIIT Jaipur(计算机科学与工程系,拉贾斯坦邦理工学院Jaipur) CSRE, IIT Bombay(印度理工学院博伊斯分校计算机科学与工程研究所)

专题命中 三维重建 :spatial understanding(abstract);分类 cs.CV

AI总结 RGFS-Net通过引入掩码图像重建任务提升遥感图像少样本分类性能,实现对未见类别的泛化和已见类别的保持,有效提高低数据下的分类准确性。

Comments Accepted at InGARSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07417 2026-01-13 cs.CV 57%

GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts

GM-MoE:基于门控机制的专家混合网络用于低光照增强

Minwen Liao, Hao Bo Dong, Xinyi Wang, Kurban Ubul, Yihua Shao, Ziyang Yan

机构 * Xinjiang University(新疆大学) Harbin University of Commerce(哈尔滨商业大学) Changchun University of Science and Technology(长春理工大学) University of Science and Technology Beijing(北京科技大学) University of Trento(特伦托大学)

专题命中 三维重建 :3D reconstruction(abstract);分类 cs.CV

AI总结 GM-MoE通过门控机制混合专家网络实现低光照图像增强,提升图像质量与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06831 2026-01-13 cs.CV 57%

SARA: Scene-Aware Reconstruction Accelerator

SARA:场景感知重建加速器

Jee Won Lee, Hansol Lim, Minhyeok Im, Dohyeon Lee, Jongseong Brad Choi

机构 * Department of Mechanical Engineering, State University of New York, Korea, Incheon, South Korea(纽约州立大学机械工程系) Department of Mechanical Engineering, State University of New York, Stony Brook, NY, United States(纽约州立大学石溪分校机械工程系) Department of Computer Science, State University of New York, Korea, Incheon, South Korea(纽约州立大学计算机科学系) Department of Computer Science, State University of New York, Stony Brook, NY, United States(纽约州立大学石溪分校计算机科学系)

专题命中 三维重建 :Gaussian Splatting(abstract);分类 cs.CV

AI总结 SARA通过几何驱动的配对选择方法,在减少配对数量的同时提升重建精度,实现高效且准确的结构从运动处理。

Comments This work has been submitted to the 2026 International Conference on Pattern Recognition (ICPR) for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16320 2026-01-13 cs.RO cs.LG 57%

PCF-Grasp: Converting Point Completion to Geometry Feature to Enhance 6-DoF Grasp

PCF-Grasp: 将点完成转换为几何特征以增强6自由度抓取

Yaofeng Cheng, Fusheng Zha, Wei Guo, Pengfei Wang, Chao Zeng, Lining Sun, Chenguang Yang

机构 * State Key Laboratory of Robotics and System at Harbin Institute of Technology(哈尔滨工业大学机器人系统国家重点实验室) Lanzhou University of Technology(兰州理工大学) Department of Computer Science(计算机科学系) University of Liverpool(利物浦大学)

专题命中 三维重建 :point cloud(abstract);分类 cs.RO

AI总结 PCF-Grasp通过将点完成转换为几何特征,提升6自由度抓取的准确性和效率,实验显示其在真实环境中的成功率比现有方法高17.8%。

Journal ref IEEE Transactions on Systems, Man, and Cybernetics: Systems ( Volume: 56, Issue: 1, January 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07139 2026-01-13 cs.CE 50%

AdaField: Generalizable Surface Pressure Modeling with Physics-Informed Pre-training and Flow-Conditioned Adaptation

AdaField:基于物理信息预训练和流条件适应的通用表面压力建模

Junhong Zou, Wei Qiu, Zhenxu Sun, Xiaomei Zhang, Zhaoxiang Zhang, Xiangyu Zhu

专题命中 三维重建 :point cloud(abstract)

AI总结 AdaField通过物理信息预训练和流条件适应,实现通用表面压力建模,适用于多种交通运输场景。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. Gaussian Splatting 13 篇

2601.06285 2026-01-13 cs.CV cs.RO 86%

NAS-GS: Noise-Aware Sonar Gaussian Splatting

NAS-GS: 噪声感知的声纳高斯点云技术

Shida Xu, Jingqi Jiang, Jonatan Scharff Willners, Sen Wang

机构 * I-X and Department of Electrical and Electronic Engineering, Imperial College London, UK(I-X 和 电气与电子工程系,帝国理工学院伦敦分校) Frontier Robotics, The National Robotarium, Edinburgh UK(前沿机器人技术,国家机器人中心,爱丁堡)

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);3D reconstruction(abstract);novel view synthesis(abstract);分类 cs.CV、cs.RO

AI总结 NAS-GS通过双向点云技术和高斯混合噪声模型,提升声纳图像的3D重建和新视角合成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02283 2026-01-13 cs.CV cs.AI 85%

GP-GS: Gaussian Processes Densification for 3D Gaussian Splatting

GP-GS:基于高斯过程的3D高斯点云密集化

Zhihao Guo, Jingxuan Su, Chenghao Qian, Shenglin Wang, Jinlong Fan, Jing Zhang, Wei Zhou, Hadi Amirpour, Yunlong Zhao, Liangxiu Han, Peng Wang

机构 * Manchester Metropolitan University(曼彻斯特 Metropolitan 大学) SECE, Peking University(SECE,北京大学) University of Leeds(利兹大学) Pengcheng Laboratory(鹏城实验室) Hangzhou Dianzi University(杭州电子科技大学) Wuhan University(武汉大学) Cardiff University(卡迪夫大学) University of Klagenfurt(克雷夫大学) Imperial College London(伦敦帝国学院)

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);3DGS(abstract);point cloud(abstract);分类 cs.CV

AI总结 GP-GS通过高斯过程实现3D高斯点云的密集化优化,提升重建质量和渲染保真度,达到1.12 dB PSNR的改进。

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19142 2026-01-13 cs.CV 85%

CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting

CLIP-GS: 通过3D高斯点划法统一视觉-语言表示

Siyu Jiao, Haoye Dong, Yuyang Yin, Zequn Jie, Yinlong Qian, Yao Zhao, Humphrey Shi, Yunchao Wei

机构 * Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学学院) National University of Singapore(新加坡国立大学) Meituan(美团) Georgia Institute of Technology(佐治亚理工学院) Picsart AI Research (PAIR)(Picsart AI研究(PAIR))

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);3DGS(abstract);point cloud(abstract);分类 cs.CV

AI总结 CLIP-GS通过3D高斯点划法统一视觉-语言表示,利用对比损失和图像投票损失提升多模态检索和分类性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21226 2026-01-13 cs.CV 84%

Frequency-Aware Gaussian Splatting Decomposition

频感知高斯散射分解

Yishai Lavi, Leo Segre, Shai Avidan

机构 * Tel Aviv University(特拉维夫大学)

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);novel view synthesis(abstract);分类 cs.CV;3D vision(comments)

AI总结 本文提出一种频感知高斯散射分解方法,通过分组和正则化提升3D重建质量与渲染效率,支持动态细节级别渲染等新功能。

Comments Accepted to the International Conference on 3D Vision (3DV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03422 2026-01-13 cs.RO cs.CV 82%

What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models

机器人中最佳的3D场景表示是什么?从几何到基础模型

Tianchen Deng, Yue Pan, Shenghai Yuan, Dong Li, Chen Wang, Mingrui Li, Long Chen, Lihua Xie, Danwei Wang, Jingchuan Wang, Javier Civera, Hesheng Wang, Weidong Chen

专题命中 Gaussian Splatting :NeRF(abstract);Gaussian Splatting(abstract);3DGS(abstract);point cloud(abstract)

AI总结 本文探讨了机器人中最佳的3D场景表示方法,对比了传统和神经表示的优劣,并展望了基础模型在机器人应用中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20031 2026-01-13 cs.RO cs.CV 81%

MG-SLAM: Structure Gaussian Splatting SLAM with Manhattan World Hypothesis

MG-SLAM:基于曼哈顿世界假设的结构高斯点撒SLAM

Shuhong Liu, Tianchen Deng, Heng Zhou, Liuzhuozheng Li, Hongyu Wang, Danwei Wang, Mingrui Li

机构 * Department of Information Science and Technology and Department of Complexity Science and Engineering, The University of Tokyo(信息科学与技术系和复杂科学与工程系,东京大学) Institute of Medical Robotics and Department of Automation, Shanghai Jiao Tong University(医疗机器人研究所和自动化系,上海交通大学) Department of Mechanical Engineering, Columbia University(机械工程系,哥伦比亚大学) School of Electrical and Electronic Engineering, Nanyang Technological University(电子与电气工程学院,南洋理工大学) Department of Computer Science, Dalian University of Technology(计算机科学系,大连理工大学)

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);分类 cs.CV、cs.RO

AI总结 MG-SLAM基于曼哈顿世界假设,通过融合线段和平面假设提升室内场景重建的几何精度和完整性,实现高斯SLAM的先进性能。

Comments IEEE Transactions on Automation Science and Engineering

Journal ref IEEE Transactions on Automation Science and Engineering 22 (2025) 17034-17049

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01235 2026-01-13 cs.CV 74%

Compensating Spatiotemporally Inconsistent Observations for Online Dynamic 3D Gaussian Splatting

补偿时空不一致的观测以实现在线动态3D高斯点散射

Youngsik Yun, Jeongmin Bae, Hyunseung Son, Seoha Kim, Hahyun Lee, Gun Bang, Youngjung Uh

机构 * Yonsei University(延世大学) Electronics and Telecommunications Research Institute(电子电信研究院)

专题命中 Gaussian Splatting :Gaussian Splatting(title);分类 cs.CV

AI总结 本文提出了一种方法,通过补偿时空不一致的观测来提升在线动态3D高斯点散射的时间一致性和渲染质量。

Comments SIGGRAPH 2025, Project page: https://bbangsik13.github.io/OR2

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07518 2026-01-13 cs.CV cs.AI 70%

Mon3tr: Monocular 3D Telepresence with Pre-built Gaussian Avatars as Amortization

Mon3tr: 单目3D远程存在与预构建高斯人偶作为记忆化

Fangyu Lin, Yingdong Hu, Zhening Liu, Yufan Zhuang, Zehong Lin, Jun Zhang

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学)

专题命中 Gaussian Splatting :Gaussian Splatting(abstract);3DGS(abstract);分类 cs.CV

AI总结 Mon3tr通过单目3DGS技术实现远程存在,利用预构建的高斯人偶降低系统复杂性,实现实时高精度3D可视化与低带宽传输。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07484 2026-01-13 cs.GR 70%

R3-RECON: Radiance-Field-Free Active Reconstruction via Renderability

R3-RECON:通过可渲染性进行无辐射场主动重建

Xiaofeng Jin, Matteo Frosi, Yiran Guo, Matteo Matteucci

专题命中 Gaussian Splatting :Gaussian Splatting(abstract);3DGS(abstract);分类 cs.GR

AI总结 R3-RECON通过无辐射场的可渲染性场实现高效主动重建,提升新视角质量和3DGS重建精度。

Comments 18 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18992 2026-01-13 cs.CV 70%

VPGS-SLAM: Voxel-based Progressive 3D Gaussian SLAM in Large-Scale Scenes

基于体素的渐进式3D高斯SLAM:用于大规模场景的VPGS-SLAM

Tianchen Deng, Wenhua Wu, Junjie He, Yue Pan, Shenghai Yuan, Danwei Wang, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(自动化与智能感知学院,上海交通大学) Key Laboratory of System Control and Information Processing, Ministry of Education(系统控制与信息处理重点实验室,教育部) Thrust of Robotics and Autonomous Systems, The Hong Kong University of Science and Technology (Guangzhou)(机器人与自主系统研究 thrust,香港科技大学(广州)) University of Bonn(波恩大学) School of Electrical and Electronic Engineering, Nanyang Technological University(电子与电气工程学院,南洋理工大学)

专题命中 Gaussian Splatting :Gaussian Splatting(abstract);3DGS(abstract);分类 cs.CV

AI总结 VPGS-SLAM提出了一种基于体素的渐进式3D高斯SLAM方法,适用于大规模室内外场景,通过多子地图实现紧凑准确的场景表示,并结合2D-3D融合跟踪和回环闭合方法提升鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05394 2026-01-13 cs.CV cs.GR cs.MM eess.IV 62%

Sketch&Patch++: Efficient Structure-Aware 3D Gaussian Representation

Sketch&Patch++: 高效的结构感知3D高斯表示

Yuang Shi, Géraldine Morin, Simone Gasparini, Wei Tsang Ooi

机构 * National University of Singapore(新加坡国立大学) IRIT - Université de Toulouse(图卢兹大学IRIT)

专题命中 Gaussian Splatting :3DGS(abstract);分类 cs.CV、cs.GR

AI总结 Sketch&Patch++提出一种结构感知的3D高斯表示方法,通过分层自适应分类框架实现高效存储和渲染,相比传统方法在PSNR、SSIM和LPIPS上均取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06479 2026-01-13 cs.CV 57%

SRFlow: A Dataset and Regularization Model for High-Resolution Facial Optical Flow via Splatting Rasterization

SRFlow:一种通过点撒射法栅格化实现高分辨率面部光流的数据库和正则化模型

JiaLin Zhang, Dong Li

机构 * School of Automation, Guangdong University of Technology(广东工业大学自动化学院)

专题命中 Gaussian Splatting :Gaussian Splatting(abstract);分类 cs.CV

AI总结 SRFlow提出高分辨率面部光流数据集和模型,通过定制正则化损失提升光流估计精度和微表情识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20377 2026-01-13 cs.CV cs.MM 57%

SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images

SmartSplat: 基于特征的智能高斯用于超高清图像的可扩展压缩

Linfei Li, Lin Zhang, Zhong Wang, Ying Shen

专题命中 Gaussian Splatting :Gaussian Splatting(abstract);分类 cs.CV

AI总结 SmartSplat通过特征感知的高斯方法实现超高清图像的高效压缩,兼顾压缩比与重建质量,展现强可扩展性。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 点云 9 篇

2601.06839 2026-01-13 cs.CV 83%

PRISM: Color-Stratified Point Cloud Sampling

PRISM:基于颜色分层的点云采样

Hansol Lim, Minhyeok Im, Jongseong Brad Choi

机构 * Department of Mechanical Engineering, State University of New York, Korea, Incheon, South Korea(纽约州立大学机械工程系) Department of Mechanical Engineering, State University of New York, Stony Brook, NY, United States(纽约州立大学石溪分校机械工程系) Department of Computer Science, State University of New York, Korea, Incheon, South Korea(纽约州立大学计算机科学系) Department of Computer Science, State University of New York, Stony Brook, NY, United States(纽约州立大学石溪分校计算机科学系)

专题命中 点云 :point cloud(title,abstract);3D reconstruction(abstract);分类 cs.CV

AI总结 PRISM通过颜色引导分层采样方法,在RGB-LiDAR点云中保留高颜色变化区域,减少视觉同质表面,提升3D重建的点云质量。

Comments This work has been submitted to the 2026 International Conference on Pattern Recognition (ICPR) for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07621 2026-01-13 cs.CG math.OC 71%

Searching point patterns in point clouds describing local topography

在描述局部地形的点云中搜索点模式

Ewa Bednarczuk, Rafał Bieńkowski, Robert Kłopotek, Jan Kryński, Krzysztof Leśsniewski, Krzysztof Rutkowski, Małgorzata Szelachowska

专题命中 点云 :point cloud(title)

AI总结 该研究提出了一种基于有限差分算子的局部描述符,用于在点云中比较和对齐结构化几何模式,并结合沃舍斯坦距离和Procrustes分析进行点分布比较和几何结构对齐。

Comments 9 pages, 2 figures, 2 tables, 6 citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07119 2026-01-13 cs.DC cs.CV 57%

SC-MII: Infrastructure LiDAR-based 3D Object Detection on Edge Devices for Split Computing with Multiple Intermediate Outputs Integration

SC-MII:基于基础设施激光雷达的边缘设备上用于分割计算的多中间输出集成的3D目标检测

Taisuke Noguchi, Takayuki Nishio, Takuya Azumi

机构 * Graduate School of Science and Engineering(科学与工程研究生学校) Saitama University(上野大学) School of Engineering(工程学院) Institute of Science Tokyo(东京科学研究所)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 SC-MII通过多激光雷达和边缘计算实现高效3D目标检测,降低延迟和能耗,提升隐私保护。

Comments 6 pages. This version includes minor lstlisting configuration adjustments for successful compilation. No changes to content or layout. Originally published at IEEE CCNC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03519 2026-01-13 cs.RO 57%

A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving

一种具有视觉提示的视觉-语言-动作模型用于越野自动驾驶

Liangdong Zhang, Yiming Nie, Haoyang Li, Fanjie Kong, Baobao Zhang, Shunxin Huang, Kai Fu, Chen Min, Liang Xiao

专题命中 点云 :spatial understanding(abstract);分类 cs.RO

AI总结 本文提出OFF-EMMA模型,通过视觉提示和COT-SC策略提升越野自动驾驶轨迹规划的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06496 2026-01-13 cs.CV 57%

3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence

3D CoCa v2:基于测试时间搜索的对比学习用于通用空间智能

Hao Tang, Ting Huang, Zeyu Zhang

机构 * School of Computer Science, Peking University(北京大学计算机学院)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 3D CoCa v2通过测试时间搜索提升3D场景描述的泛化能力,实现对比学习与描述生成的统一,提高鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06465 2026-01-13 eess.IV cs.CV cs.MM 57%

R$^3$D: Regional-guided Residual Radar Diffusion

R$^3$D: 基于区域引导的残差雷达扩散

Hao Li, Xinqi Liu, Yaoqing Jin

机构 * University of Arizona(亚利桑那大学) University of Hong Kong(香港大学) University of Stuttgart(斯图加特大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 R3D通过区域引导残差雷达扩散框架,整合残差建模与sigma自适应引导,提升雷达点云质量,优于现有方法。

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06464 2026-01-13 cs.CV 57%

On the Adversarial Robustness of 3D Large Vision-Language Models

关于3D大视觉-语言模型的对抗鲁棒性

Chao Liu, Ngai-Man Cheung

机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)

专题命中 点云 :3D vision(abstract);分类 cs.CV

AI总结 本文研究了3D大视觉-语言模型的对抗鲁棒性,提出两种攻击策略评估其鲁棒性,发现其在无目标攻击下较脆弱,但在有目标攻击中更稳健。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19290 2026-01-13 cs.CV 57%

TRASE: Tracking-free 4D Segmentation and Editing

TRASE:无需跟踪的4D分割与编辑

Yun-Jin Li, Mariia Gladkova, Yan Xia, Daniel Cremers

机构 * TU Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 点云 :3D reconstruction(abstract);分类 cs.CV

AI总结 TRASE通过弱监督学习实现无需跟踪的4D分割,利用对比学习和聚类技术实现动态场景的高效分割与交互编辑。

Comments Accepted to 3DV 2026. Project page https://yunjinli.github.io/project-sadg

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07099 2026-01-13 eess.SP 50%

Autofocus Method for Human-Body Imaging under Respiratory Motion Using Synthetic Aperture Radar

用于呼吸运动下人体成像的自动聚焦方法:合成孔径雷达

Masaya Kato, Takuya Sakamoto

专题命中 点云 :point cloud(abstract)

AI总结 本研究提出一种用于呼吸运动下人体合成孔径雷达成像的自动聚焦方法,通过分离回波和估计相位误差提升图像质量。

Comments 8 pages, 7 figures, and 3 tables. This work is going to be submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 新视角合成 1 篇

2601.04382 2026-01-13 cs.GR cs.CV 62%

Radiant Foam Rendering on a Graph Processor

图处理器上的辐射泡沫渲染

Zulkhuu Tuya, Ignacio Alzugaray, Nicholas Fry, Andrew J. Davison

机构 * Imperial College London(帝国理工学院伦敦分校)

专题命中 新视角合成 :NeRF(abstract);分类 cs.CV、cs.GR

AI总结 在Graphcore Mk2 IPU上实现辐射泡沫体渲染的高效分布式渲染器,通过分层路由和本地SRAM实现高吞吐量和高质量渲染。

Comments 24 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 空间理解 2 篇

2503.22976 2026-01-13 cs.CV 57%

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

从二维空间到三维空间:教视觉语言模型感知和推理三维世界

Jiahui Zhang, Yurui Chen, Yanpeng Zhou, Yueming Xu, Ze Huang, Jilin Mei, Junhui Chen, Yu-Jie Yuan, Xinyue Cai, Guowei Huang, Xingyue Quan, Hang Xu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 本文提出了一种基于2D空间数据生成和标注的管道,构建了SPAR-7M数据集和SPAR-Bench基准,通过训练和微调提升视觉语言模型在三维空间感知和推理中的性能。

Comments Project page: https://logosroboticsgroup.github.io/SPAR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06781 2026-01-13 cs.HC cs.AI cs.CV 57%

AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs

AutoTour:基于智能手机和LLMs的自动照片导览系统

Huatao Xu, Zihe Liu, Zilin Zeng, Baichuan Li, Mo Li

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 AutoTour利用智能手机和LLMs自动为用户照片生成细粒度地标注释和描述,实现可扩展且上下文感知的交互式探索体验。

Comments 21

详情

展开后加载摘要…

URL PDF HTML 收藏