arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 8673 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 机器人数据与评测 8673 篇

2512.19551 2025-12-23 cs.AI 57%

Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen Scenarios

迈向闭环式具身共情进化:探索以LLM为中心的终身共情动作生成在未见场景中的能力

Jiawen Wang, Jingjing Wang Tianyang Chen, Min Zhang, Guodong Zhou

专题命中 机器人数据与评测 :embodied agent(abstract);分类 cs.AI

AI总结 本文提出L^2-EMG任务,旨在通过情感解耦和场景适应挑战提升LLM在未见场景中的共情动作生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19234 2025-12-23 cs.AI 57%

DeliveryBench: Can Agents Earn Profit in Real World?

DeliveryBench: 代理在现实世界中能否获利?

Lingjun Mao, Jiawei Ren, Kun Zhou, Jixuan Chen, Ziqiao Ma, Lianhui Qin

机构 * University of California, San Diego(加州大学圣地亚哥分校) University of Michigan(密歇根大学)

专题命中 机器人数据与评测 :embodied agent(abstract);分类 cs.AI

AI总结 DeliveryBench通过模拟现实世界中的食品配送任务,评估基于VLM的具身代理在长期规划和约束管理方面的性能,发现其与人类存在显著差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19088 2025-12-23 cs.CV 57%

Retrieving Objects from 3D Scenes with Box-Guided Open-Vocabulary Instance Segmentation

通过框引导的开放词汇实例分割从3D场景中检索物体

Khanh Nguyen, Dasith de Silva Edirimuni, Ghulam Mubashar Hassan, Ajmal Mian

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 本文提出了一种基于2D开放词汇检测器引导的3D实例分割方法,用于从RGB图像中快速准确检索稀有物体实例。

Comments Accepted to AAAI 2026 Workshop on New Frontiers in Information Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14501 2025-12-23 cs.CV 57%

Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey

反向馈送3D重建与视角合成的进展:综述

Jiahui Zhang, Yuelei Li, Anpei Chen, Muyu Xu, Kunhao Liu, Jianyuan Wang, Xiao-Xiao Long, Hanxue Liang, Zexiang Xu, Hao Su, Christian Theobalt, Christian Rupprecht, Andrea Vedaldi, Kaichen Zhou, Hanspeter Pfister, Paul Pu Liang, Shijian Lu, Fangneng Zhan

机构 * NTU(国立台湾大学) Caltech(加州理工学院) Westlake University(西湖大学) UCSD(加州大学圣地亚哥分校) University of Oxford(牛津大学) Nanjing University(南京大学) HKU(香港大学) University of Cambridge(剑桥大学) Hillbot MPI for Informatics(信息研究所) Harvard University(哈佛大学) MIT(麻省理工学院)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 本文综述了反向馈送方法在3D重建与视角合成中的进展,涵盖多种表示架构及应用,探讨了关键任务和未来研究方向。

Comments A project page associated with this survey is available at https://fnzhan.com/projects/Feed-Forward-3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18456 2025-12-23 cs.CR cs.AI 57%

SoK: Understanding (New) Security Issues Across AI4Code Use Cases

SoK:理解(新的)AI4Code使用案例中的安全问题

Qilong Wu, Taoran Li, Tianyang Zhou, Varun Chandrasekaran

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI

AI总结 本文通过调查AI4Code在三个核心应用中的安全问题,指出基准测试偏见、数据泄露和对抗鲁棒性不足等挑战,并提出安全默认实践、全面检测基准和翻译增强语言等改进方向。

Comments 39 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03991 2025-12-22 cs.HC cs.RO 57%

When to Say "Hi" -- Learn to Open a Conversation with an in-the-wild Dataset

何时说‘你好’——利用真实场景数据集学习开启对话

Michael Schiffmann, Felix Struth, Sabina Jeschke, Anja Richert

机构 * Cologne Cobots Lab, TH Köln - University of Applied Sciences(科隆机器人实验室,TH Köln应用科学大学) KI Park e.V. & FAU - Friedrich-Alexander University of Erlangen-Nuremberg(KI Park协会与弗赖堡-艾尔朗根-纽伦堡大学)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.RO

AI总结 本文提出IIS系统,通过真实场景数据集训练,实现基于用户身体语言的对话开启识别与判断。

Comments 6 pages, 3 figures, 5 tables. This paper has been accepted for publication at IEEE ROMAN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06278 2025-12-22 cs.RO 57%

Mitigating Undesired Conditions in Flexible Production with Product-Process-Resource Asset Knowledge Graphs

通过产品-过程-资源资产知识图谱缓解柔性生产中的不良状况

Petr Novak, Stefan Biffl, Marek Obitko, Petr Kadera

机构 * Czech Institute of Informatics, Robotics and Cybernetics(捷克信息学、机器人学与自动控制研究所) Czech Technical University in Prague(布拉格捷克技术大学) TU Wien -- Institute of Information Systems Engineering, Faculty of Informatics(维也纳技术大学——信息系统工程研究所,信息学学院)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 本文提出PPR-AKG模型,通过结合语义技术和大语言模型,解决柔性生产中的不良状况问题,提升资源分配效率和生产质量。

Comments Originally published online by CEUR Workshop Proceedings (CEUR-WS.org, ISSN 1613-0073) within ISWC 2025 Companion Volume. Available online: https://ceur-ws.org/Vol-4085/ and https://ceur-ws.org/Vol-4085/paper9.pdf

Journal ref CEUR Workshop Proceedings (CEUR-WS.org), ISSN 1613-0073. ISWC 2025 Companion Volume. Available online: https://ceur-ws.org/Vol-4085/ and https://ceur-ws.org/Vol-4085/paper9.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17349 2025-12-22 cs.RO 57%

Flying in Clutter on Monocular RGB by Learning in 3D Radiance Fields with Domain Adaptation

在仅用单目RGB图像中飞行的 clutter 问题通过学习3D辐射场并进行域适应

Xijie Huang, Jinhan Li, Tianyue Wu, Xin Zhou, Zhichao Han, Fei Gao

机构 * State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China(浙江大学工业控制技术状态重点实验室) Differential Robotics, Hangzhou 311121, China(差异机器人技术)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.RO

AI总结 本文提出通过学习3D辐射场并结合域适应,使飞行机器人在仅使用单目RGB图像的情况下实现鲁棒的零样本迁移,以应对 clutter 和光照变化的挑战。

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11362 2025-12-22 cs.RO 57%

An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges

视觉-语言-动作模型的解剖:从模块到里程碑与挑战

Chao Xu, Suyu Zhang, Yang Liu, Baigui Sun, Weihong Chen, Bo Xu, Qi Liu, Juncheng Wang, Shujun Wang, Shan Luo, Jan Peters, Athanasios V. Vasilakos, Stefanos Zafeiriou, Jiankang Deng

机构 * IROOTECH TECHNOLOGY(IROOTECH技术公司) Wolf 1069 b Lab, Sany Group(Sany集团沃尔夫1069b实验室) Department of Engineering, King’s College London(伦敦国王学院工程系) Hong Kong Polytechnic University(香港理工大学) Computer Science Department of the Technische Universität Darmstadt(德累斯顿技术大学计算机科学系) Department of ICT and Center for AI Research, University of Agder (UiA)(阿格德大学信息与通信技术系及人工智能研究中心) Department of Computing, Imperial College London(伦敦帝国理工学院计算系)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.RO

AI总结 本文系统分析了视觉-语言-动作模型的核心挑战,从模块构建到里程碑发展,为研究者提供结构化指南和未来研究方向。

Comments project page: https://suyuz1.github.io/VLA-Survey-Anatomy/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00780 2025-12-22 cs.RO cs.SY eess.SY 57%

Accelerating Hybrid Model Predictive Control using Warm-Started Generalized Benders Decomposition

利用预热的广义邦德斯分解加速混合模型预测控制

Xuan Lin

机构 * Independent Researcher(独立研究者) Department of Mechanical Engineering, University of California, Los Angeles(加州大学洛杉矶分校机械工程系) School of Mechanical Engineering, Georgia Institute of Technology(佐治亚理工学院机械工程学院)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 本文提出一种基于广义邦德斯分解的混合MPC算法,通过预热切割平面提升求解速度,适用于机器人控制任务。

Comments Significantly revised to emphasize theoretical bounds. The heuristic Master Problem algorithm and GCS tightening experiments are preserved in v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16493 2025-12-19 cs.CV 57%

YOLO11-4K: An Efficient Architecture for Real-Time Small Object Detection in 4K Panoramic Images

YOLO11-4K: 一种高效的实时小目标检测架构用于4K全景图像

Huma Hafeez, Matthew Garratt, Jo Plested, Sankaran Iyer, Arcot Sowmya

机构 * School of Engineering \& Technology, University of New South Wales, Canberra, Australia School of Systems \& Computing, University of New South Wales, Canberra, Australia School of Computer Science \& Engineering, University of New South Wales, Sydney, Australia

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.CV

AI总结 YOLO11-4K通过高效架构实现4K全景图像中小目标的实时高精度检测,相比YOLO11在精度和速度上均有显著提升。

Comments Conference paper just submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16213 2025-12-19 cs.CV math.DG 57%

Enhanced 3D Shape Analysis via Information Geometry

通过信息几何增强的3D形状分析

Amit Vishwakarma, K. S. Subrahamanian Moosath

机构 * Indian Institute of Space Science and Technology(印度空间科学与技术研究所)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 本文提出基于信息几何的3D点云分析方法,通过高斯混合模型和改进的KL散度实现稳定且准确的形状比较。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12793 2025-12-19 cs.RO 57%

VLG-Loc: Vision-Language Global Localization from Labeled Footprint Maps

VLG-Loc: 从标注足迹地图实现视觉-语言全局定位

Mizuho Aoki, Kohei Honda, Yasuhiro Yoshimura, Takeshi Ishita, Ryo Yonetani

机构 * The Department of Mechanical Systems Engineering, Nagoya University(名古屋大学机械系统工程系) CyberAgent AI Lab(CyberAgent AI实验室)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 VLG-Loc通过视觉-语言模型实现从标注足迹地图的全局定位,结合蒙特卡洛定位框架和概率融合提升环境变化下的鲁棒性。

Comments v2: Updated the citation of SparseLoc from an arXiv preprint to its published version in the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07392 2025-12-19 cs.CL cs.AI 57%

Voice-Interactive Surgical Agent for Multimodal Patient Data Control

多模态患者数据控制的语音交互手术代理

Hyeryun Park, Byung Mo Gu, Jun Hee Lee, Byeong Hyeon Choi, Sekeun Kim, Hyun Koo Kim, Kyungsang Kim

机构 * Image Guided Precision Cancer Surgery Institute, College of Medicine, Korea University(影像引导精准癌症手术研究所,韩国大学医学院) Department of Radiology, Massachusetts General Hospital(放射科,麻省总医院) Department of Thoracic and Cardiovascular Surgery, Korea University Guro Hospital, College of Medicine, Korea University(胸外科和心血管外科,韩国大学Guro医院,韩国大学医学院) Department of Biomedical Sciences, College of Medicine, Korea University(生物医学科学系,韩国大学医学院)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.AI

AI总结 本文提出一种基于大型语言模型的语音交互手术代理,用于多模态患者数据的高效控制与处理,通过分层多代理框架提升手术流程中的数据操作效率与鲁棒性。

Comments 14 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15483 2025-12-18 cs.LG 57%

Multi-stage Bayesian optimisation for dynamic decision-making in self-driving labs

多阶段贝叶斯优化用于自动驾驶实验室中的动态决策

Luca Torresi, Pascal Friederich

机构 * Institute of Nanotechnology, Karlsruhe Institute of Technology(纳米技术研究所,卡尔斯鲁厄技术大学) Institute of Anthropomatics and Robotics, Karlsruhe Institute of Technology(人机动力学与机器人研究所,卡尔斯鲁厄技术大学)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.LG

AI总结 本文提出多阶段贝叶斯优化方法,通过引入代理测量提升自动驾驶实验室中动态决策的效率和优化效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02179 2025-12-18 cs.NE cs.AI 57%

Towards Robust and Accurate Myoelectric Controller Design based on Multi-objective Optimization using Evolutionary Computation

基于多目标优化的鲁棒且准确的肌电控制器设计:利用进化计算

Ahmed Aqeel Shaikh, Anand Kumar Mukhopadhyay, Soumyajit Poddar, Suman Samui

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.AI

AI总结 本文提出基于多目标优化和进化计算的肌电控制器设计方法,通过优化分类器参数提升控制器的鲁棒性和准确性。

Comments This is the updated paper

Journal ref IEEE Sensors Journal ( Volume: 24, Issue: 5, 01 March 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14434 2025-12-17 cs.RO 57%

Geometric Parameter Optimization of a Novel 3-(PP(2-(UPS))) Redundant Parallel Mechanism based on Workspace Determination

一种基于工作空间确定的新型3-(PP(2-(UPS)))冗余并联机构几何参数优化

Quan Yuan, Daqian Cao, Weibang Bai

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 本文提出了一种新型冗余并联机构,通过分析关键几何参数对工作空间的影响,优化其几何参数以提升工作空间性能。

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14331 2025-12-17 cs.RO 57%

ARCADE: Adaptive Robot Control with Online Changepoint-Aware Bayesian Dynamics Learning

ARCADE: 基于在线变化点感知的自适应机器人控制

Rishabh Dev Yadav, Avirup Das, Hongyu Song, Samuel Kaski, Wei Pan

机构 * Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系) Department of Computer Science, Aalto University(阿alto大学计算机科学系)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 ARCADE通过在线变化点感知机制实现机器人动态的自适应控制,提升预测精度和恢复能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.01519 2025-12-17 cs.AI 57%

A Generalised Approach for Encoding and Reasoning with Qualitative Theories in Answer Set Programming

一种用于答案集编程中定性理论编码与推理的通用方法

George Baryannis, Ilias Tachmazidis, Sotiris Batsakis, Grigoris Antoniou, Mario Alviano, Emmanuel Papadakis

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI

AI总结 本文提出了一种基于答案集编程的通用方法,用于处理需要结合定性与非定性推理的问题,并通过实验验证了其有效性。

Comments Paper presented at the 36th International Conference on Logic Programming (ICLP 2020), University Of Calabria, Rende (CS), Italy, September 2020, 18 pages, 3 figures

Journal ref Theory and Practice of Logic Programming 20 (2020) 687-702

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12395 2025-12-16 cs.CV 57%

ArtGen: Conditional Generative Modeling of Articulated Objects in Arbitrary Part-Level States

ArtGen: 条件生成建模任意部分级状态的关节物体

Haowen Wang, Xiaoping Yuan, Fugang Zhang, Rui Jian, Yuanwei Zhu, Xiuquan Qiao, Yakun Huang

机构 * Anhui University(安徽大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 ArtGen通过条件扩散框架生成任意部分级状态的关节3D物体,采用跨状态采样和思维链推理模块,提升运动学一致性与结构先验推断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12377 2025-12-16 cs.RO 57%

INDOOR-LiDAR: Bridging Simulation and Reality for Robot-Centric 360 degree Indoor LiDAR Perception -- A Robot-Centric Hybrid Dataset

INDOOR-LiDAR: 联接仿真与现实以提升机器人中心的360度室内LiDAR感知 -- 一个机器人中心的混合数据集

Haichuan Li, Changda Tian, Panos Trahanias, Tomi Westerlund

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 INDOOR-LIDAR通过融合仿真与真实数据,为机器人感知研究提供了一个可扩展、真实且可重复的混合数据集,用于提升复杂室内环境中的感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12919 2025-12-16 cs.CV 57%

CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation

CoordAR:通过自回归坐标图生成实现新型物体的单参考6D姿态估计

Dexin Zuo, Ang Li, Wei Wang, Wenxian Yu, Danping Zou

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 CoordAR通过自回归坐标图生成方法,实现对新型物体的单参考6D姿态估计,提升了对称性和遮挡等挑战的鲁棒性。

Comments 7 pages, accepted by AAAI 2026 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10734 2025-12-15 cs.CL cs.AI 57%

Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation

文本数据偏差检测与缓解——一个可扩展的流程及其实验评估

Rebekka Görge, Sujan Sai Gannamaneni, Tabea Naeven, Hammam Abdelwahab, Héctor Allende-Cid, Armin B. Cremers, Lennard Helmer, Michael Mock, Anna Schmitz, Songkai Xue, Elif Yildirir, Maximilian Poretschkin, Stefan Wrobel

机构 * organization= Fraunhofer Institute for Intelligent Analysis organization= Trustworthiness Theory, Technology \& Engineering Lab, Huawei Technologies Co., Ltd , country= China organization= Lamarr Institute for Machine Learning organization= B-IT Emeritus Research Group AI Foundations, University of Bonn , country= Germany

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI

AI总结 本文提出了一种可扩展的数据偏见检测与缓解流程,通过四个组件减少文本数据中的表示偏见和刻板印象,并通过实验评估展示其对模型偏见减少的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11125 2025-12-15 cs.RO cs.SY eess.SY 57%

Design and Experimental Validation of Closed-Form CBF-Based Safe Control for Stewart Platform Under Multiple Constraints

六自由度平台在多重约束下的安全控制闭环形式设计与实验验证

Benedictus C. G. Cinun, Tua A. Tamba, Immanuel R. Santjoko, Xiaofeng Wang, Michael A. Gunarso, Bin Hu

机构 * Department of Electrical and Computer Engineering at University of Houston(德克萨斯大学休斯顿分校电气与计算机工程系) Department of Electrical Engineering, Faculty of Engineering Technology, Parahyangan Catholic University(帕拉哈扬天主教大学工程技术学院电气工程系) Department of Electrical Engineering, University of South Carolina(南卡罗来纳大学电气工程系) Department of Engineering Technology, Electrical and Computer Engineering, University of Houston(德克萨斯大学休斯顿分校工程技术系电气与计算机工程系)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO

AI总结 本文提出了一种基于闭环形式的CBF方法,用于六自由度平台在多重约束下的安全控制,通过显式控制律实现高效实时控制,减少计算量并验证了安全性。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10628 2025-12-12 cs.CV 57%

K-Track: Kalman-Enhanced Tracking for Accelerating Deep Point Trackers on Edge Devices

K-Track: Kalman增强的跟踪用于加速边缘设备上的深度点跟踪器

Bishoy Galoaa, Pau Closas, Sarah Ostadabbas

机构 * Northeastern University(东北大学)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 K-Track通过结合深度学习与卡尔曼滤波,实现边缘设备上点跟踪的高效部署,提升速度的同时保持高精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10352 2025-12-12 cs.CV 57%

Topology-Agnostic Animal Motion Generation from Text Prompt

基于文本提示的拓扑无关动物运动生成

Keyi Chen, Mingze Sun, Zhenyu Liu, Zhangquan Chen, Ruqi Huang

机构 * Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 本文提出OmniZoo数据集和通用自回归框架,实现基于文本提示的任意拓扑动物运动生成及跨物种风格迁移。

Comments 10 pages, 7 figures.Conference submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10293 2025-12-12 cs.CV 57%

Physically Aware 360$^\circ$ View Generation from a Single Image using Disentangled Scene Embeddings

基于单张图像的物理感知360度视图生成

Karthikeya KV, Narendra Bandaru

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.CV

AI总结 Disentangled360通过结合方向解耦体积渲染和单图像360°视图合成,实现了医学影像和自然场景重建中的高效真实感视图生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09619 2025-12-11 cs.RO 57%

GLaD: Geometric Latent Distillation for Vision-Language-Action Models

GLaD:面向视觉-语言-动作模型的几何潜在蒸馏

Minghao Guo, Meng Cao, Jiachen Tao, Rongtao Xu, Yan Yan, Xiaodan Liang, Ivan Laptev, Xiaojun Chang

机构 * MBZUAI University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.RO

AI总结 GLaD通过引入几何意识的预训练机制,提升了视觉-语言-动作模型的空间推理和策略泛化能力,无需依赖深度传感器或3D标注。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09995 2025-12-11 cs.CV 57%

PlayerOne: Egocentric World Simulator

PlayerOne:第一人称真实世界模拟器

Yuanpeng Tu, Hao Luo, Xi Chen, Xiang Bai, Fan Wang, Hengshuang Zhao

机构 * HKU(香港大学) DAMO Academy, Alibaba Group(阿里云达摩院) Hupan Lab(华盘实验室) HUST(华中科技大学)

专题命中 机器人数据与评测 :world model(abstract);分类 cs.CV

AI总结 PlayerOne通过第一人称真实世界模拟器实现了沉浸式探索,通过粗到细的训练流程和部分解耦运动注入方案,实现了对人类运动的精确控制和多样化场景的一致建模。

Comments Project page: https://playerone-hku.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04734 2025-12-10 cs.CV 57%

MT-Depth: Multi-task Instance feature analysis for the Depth Completion

MT-Depth:多任务实例特征分析用于深度补全

Abdul Haseeb Nizamani, Dandi Zhou, Xinhai Sun

机构 * Saisuode (Shanghai) Intelligent Technology Co., Ltd. (Synthoid.ai)(上海思速德智能科技有限公司(Synthoid.ai))

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV

AI总结 MT-Depth通过实例感知的多任务实例特征分析,提升深度补全的精度,尤其在物体边界、遮挡和细长结构区域表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏