Learning Object Localization and 6D Pose Estimation from Simulation and Weakly Labeled Real Images
专题命中 机器人数据与评测 :robotics(abstract);manipulation(abstract);robotic(abstract);分类 cs.CV
视觉与机器人
机器人、具身智能、机器人学习、操作、导航和具身世界模型。
专题命中 机器人数据与评测 :robotics(abstract);manipulation(abstract);robotic(abstract);分类 cs.CV
专题命中 机器人数据与评测 :navigation(abstract);world model(abstract);robotic(abstract);分类 cs.CV
Comments The project web page at http://vcl.itn.liu.se/publications/2017/TKWU17/ contains a version of the paper with high-resolution images as well as additional material
专题命中 机器人数据与评测 :robotics(abstract);manipulation(abstract);robotic(abstract);分类 cs.CV
专题命中 机器人数据与评测 :robotics(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO
Comments 8 pages, 10 figures
Journal ref IROS 2016
专题命中 机器人数据与评测 :robotics(abstract);navigation(abstract);robotic(abstract);分类 cs.RO
Journal ref Progrès en urologie : journal de l'Association française d'urologie et de la Société française d'urologie 16, 2 (2006) 112-20
机构 * Korea Institute of Machinery & Materials (KIMM)(韩国机械材料研究院) ; Gwangju Institute of Science and Technology (GIST)(全州科学技术院) ; Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 机器人数据与评测 :robotics(abstract,comments);robotic(abstract);分类 cs.RO、cs.AI、cs.CV
Comments Accepted at IEEE Robotics and Automation Letters (RA-L). Project Websites: https://sites.google.com/view/graspclutter6d
机构 * School of Informatics, Computing and Engineering, Indiana University(信息学、计算与工程学院,印第安纳大学) ; Nuro, Inc.(Nuro公司)
专题命中 机器人数据与评测 :robotics(abstract,journal_ref);robotic(abstract);分类 cs.RO、cs.AI、cs.LG
Journal ref IEEE ROBOTICS AND AUTOMATION LETTERS (RA-L) 2025
专题命中 机器人数据与评测 :manipulation(abstract);robot policy(abstract);分类 cs.RO、cs.AI、cs.LG;robot learning(comments)
Comments 8th Conference on Robot Learning (CoRL 2024), Munich, Germany. Project website: https://transic-robot.github.io/
专题命中 机器人数据与评测 :robotics(abstract,comments);robotic(abstract);分类 cs.RO、cs.AI、cs.LG
Comments 14 pages, 14 figures, accepted to Transactions on Robotics
专题命中 机器人数据与评测 :robotics(abstract,comments);manipulation(abstract);分类 cs.RO、cs.AI、cs.LG
Comments 4 pages, 4 figures, camera-ready manuscript, accepted to the "Causality for Robotics: Answering the Question of Why" workshop at the 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
专题命中 机器人数据与评测 :manipulation(abstract);robotic(abstract);分类 cs.RO、cs.AI、cs.LG;robot learning(comments)
Comments Accepted to Conference on Robot Learning (CoRL) 2023
专题命中 机器人数据与评测 :robotics(abstract,comments);manipulation(abstract);robotic(abstract)
Comments 20 pages with 22 figures. Submitted to IEEE Transactions on Robotics (T-RO). The supplemental video is available publicly at https://youtu.be/M_KI_cM1-rI
专题命中 机器人数据与评测 :robotics(abstract,journal_ref);manipulation(abstract);分类 cs.RO、cs.AI、cs.LG
Comments First two authors contributed equally. arXiv admin note: text overlap with arXiv:1803.00967
Journal ref The International Journal of Robotics Research (IJRR), 2020
专题命中 机器人数据与评测 :robotics(abstract,comments);robotic(abstract);分类 cs.RO、cs.CV、cs.LG
Comments To be presented at Robotics: Science and Systems 2020
专题命中 机器人数据与评测 :manipulation(abstract);robotic(abstract);分类 cs.RO、cs.CV、cs.LG;robotics(comments)
Comments Accepted to IEEE Robotics and Automation Letters and selected by IROS'18 Program Committee for presentation at the Conference
专题命中 机器人数据与评测 :navigation(abstract);robotic(abstract);分类 cs.RO、cs.CV、cs.LG;robotics(comments)
Comments To appear at Robotics: Science and Systems Conference (R:SS), 2017. Supplementary video: https://www.youtube.com/watch?v=nXBWmzFrj5s
VT-MUSE:面向操作任务的多模态统一时序视觉-触觉表示学习
机构 * Shanghai Jiao Tong University(上海交通大学) ; Xense Robotics(Xense机器人公司) ; BUPT(北京邮电大学)
专题命中 机器人数据与评测 :manipulation(title);分类 cs.RO、cs.CV
AI总结 本研究提出VT-MUSE视觉-触觉多模态统一时序表示学习框架,通过两阶段学习解决现有方法的跨模态依赖捕获与时序接触演化忽略问题,在仿真和真实操作任务中性能优于基线
带不确定时变奖励的定向越野问题:面向日常服务机器人的框架与基准
专题命中 机器人数据与评测 :robotics(title);分类 cs.RO、cs.AI
AI总结 针对日常服务机器人的带不确定时变奖励的定向越野问题,本文提出相关框架与基准,采用三种不同规划器并验证了长时域在线适应规划的有效性。
Comments 11 pages, 6 figures
DA-Fusion:基于可变形注意力的RGB-D融合Transformer用于未见物体实例分割
机构 * Interdisciplinary Program in AI, Seoul National University(首尔国立大学人工智能跨学科项目) ; Artificial Intelligence Institute, Seoul National University(首尔国立大学人工智能研究所) ; Department of Computer Science, Seoul National University(首尔国立大学计算机科学系)
专题命中 机器人数据与评测 :manipulation(abstract);robotic(abstract);robotics(comments,journal_ref);分类 cs.AI、cs.CV
AI总结 针对物流自动化中未见物体分割难题,提出基于可变形注意力的RGB-D融合Transformer(DA-Fusion),结合RGB与深度数据优势提升分割精度,引入OCBD数据集,实验证明其性能优于现有方法,适用于现实物流任务。
Comments 7 pages, 5 figures. Published in the Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA 2025)
Journal ref 2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 7490-7496
面向机器人视觉-语言模型的具有推理能力的空间轨迹
专题命中 机器人数据与评测 :robotics(title);分类 cs.RO、cs.CV
AI总结 本文提出RoboTracer,一种3D感知的视觉语言模型,通过空间编码器和回归监督解码器提升空间推理能力,并结合强化学习训练实现多步度量推理,最终在复杂场景中实现高效空间轨迹生成。
Comments Accepted to ECCV 2026. Project page: https://zhoues.github.io/RoboTracer
NarrativeWorldBench:面向长程共创音频剧的前沿饱和基准与潜在世界模型
机构 * University of California, Santa Barbara(加州大学圣塔芭芭拉分校) ; Pocket FM
专题命中 机器人数据与评测 :world model(title);分类 cs.AI、cs.LG
AI总结 提出NarrativeWorldBench基准,在九种叙事结构指标上评估21个模型,并引入N-VSSM变分状态空间模型,通过Mamba-2骨干和事件条件后验在200集以上维持结构化潜在状态,在长弧一致性和可控性上超越Claude Opus 4.5。
Comments 10 pages. Accepted to the ICML 2026 Workshops on High-dimensional Learning Dynamics (HiLD) and Culture x AI
HalluWorld: 一个用于通过参考世界模型控制幻觉的基准
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Patronus AI ; Independent Researcher(独立研究者) ; Stanford University(斯坦福大学) ; The Ohio State University(俄亥俄州立大学) ; DegenAI Labs(DegenAI实验室)
专题命中 机器人数据与评测 :world model(title);分类 cs.AI、cs.LG
AI总结 本文提出HalluWorld基准,通过显式参考世界模型研究语言模型的幻觉问题,发现不同任务中幻觉表现不一致,表明幻觉源于多种失败模式而非单一能力。
Comments HalluWorld benchmark (code and data) at github.com/DegenAI-Labs/HalluWorld
REI-Bench: 体感代理能否在任务规划中理解模糊的人类指令?
机构 * MARS Lab, School of Mechanical and Aerospace Engineering(MARS实验室,机械与航空航天工程学院)
专题命中 机器人数据与评测 :embodied agent(title);分类 cs.RO、cs.AI
AI总结 本文研究模糊参照表达对LLM任务规划的影响,提出REI-Bench基准,发现模糊性会显著降低机器人规划性能,提出任务导向上下文认知方法提升表现。
Comments Accepted at ICLR 2026
K2MUSE:一种涵盖任务和采集变异性的下肢多模态行走数据集,用于康复机器人
机构 * State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences(机器人与智能系统国家重点实验室,沈阳自动化研究所,中国科学院) ; University of Chinese Academy of Sciences(中国科学院大学) ; College of Artificial Intelligence, Tianjin Key Laboratory of Intelligent Robotics, Nankai University(人工智能学院,天津智能机器人重点实验室,南开大学) ; School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学)
专题命中 机器人数据与评测 :robotics(title);分类 cs.RO、cs.AI
AI总结 K2MUSE数据集通过多模态数据揭示下肢康复机器人与生物力学信息的关系,为康复机器人开发提供大规模步态样本和数据驱动方法支持。
Comments 34 pages, 30 figures,7 tables
具身AI数据集中的语言多样性有限
机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) ; Institute of Computer Science, University of Tartu(塔尔图大学计算机科学学院) ; Department of Mechanical Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校机械工程系) ; University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
专题命中 机器人数据与评测 :embodied AI(title);分类 cs.RO、cs.AI
AI总结 本文分析了多个广泛使用的视觉-语言-动作数据集的语言特性,发现其指令存在高度重复和结构单一的问题,旨在推动更详细的 dataset 报告和语言覆盖的扩展策略。
Comments Accepted to ACL 2026 (Main Conference)
机器人中的视觉-语言-动作:数据集、基准和数据引擎的综述
机构 * University of Maryland, College Park(马里兰大学学院公园分校) ; University of Utah(犹他大学) ; Northeastern University(东北大学) ; University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
专题命中 机器人数据与评测 :robotics(title);分类 cs.RO、cs.AI
AI总结 本文综述了视觉-语言-动作研究中数据基础设施的关键挑战,指出未来进展依赖于数据引擎与评估协议的协同设计,揭示了数据集、基准和数据引擎的四大开放挑战。
Comments This is a survey paper. The survey is already accepted by TMLR after peer-review. The OpenReview link is here: https://openreview.net/forum?id=tAaWFpvnmm
潜在偏置对齐用于高保真扩散倒置在现实世界图像重建与操控中的应用
机构 * Department of Electrical and Electronic Engineering, Southern University of Science and Technology(南方科技大学电子与电气工程系) ; Undergraduate School of Artificial Intelligence, Shenzhen Polytechnic University(深圳理工大学人工智能本科生学院) ; Huawei Technologies Co., Ltd.(华为技术有限公司) ; Pengcheng Lab(鹏城实验室)
专题命中 机器人数据与评测 :manipulation(title);分类 cs.AI、cs.CV
AI总结 本文提出Latent Bias Optimization和Image Latent Boosting方法,解决扩散模型与现实场景衔接中的重建质量与鲁棒性问题,提升图像编辑和稀有概念生成性能。
非语言实时人机交互在受限机器人环境中
机构 * National University of Science ; Princeton University, Princeton NJ 08544, USA, , WWW home page: http://users/ iekeland/web/welcome.html Universit\' e de Paris-Sud, Laboratoire d'Analyse Num\' e rique, B\ a timent 425, F-91405 Orsay Cedex, France
专题命中 机器人数据与评测 :robotic(title);分类 cs.AI、cs.CV
AI总结 本文提出首个实时生成非语言人机交互的框架,通过对比合成与真实数据发现时间一致性对实际性能的影响,揭示了人类与AI运动间的统计差异。
通过风格识别的循环一致生成对抗网络实现仿真到现实迁移:通过视觉领域适应实现机器人机械臂的零样本部署
机构 * Institute for Research in Technology (IIT)(技术研究 institute) ; ICAI School of Engineering(工程学院) ; Comillas Pontifical University(宗座大学)
专题命中 机器人数据与评测 :robotic(title);分类 cs.RO、cs.AI
AI总结 本文提出基于风格识别循环一致生成对抗网络的领域适应方法,实现机器人机械臂在现实环境中的零样本部署,通过视觉领域适应提升仿真到现实迁移的效率和准确性。
Journal ref Engineering Applications of Artificial Intelligence, volume 159, published Jan.2026
TwoHead-SwinFPN:一种统一的深度学习架构,用于身份文档中的合成操纵、检测和定位
机构 * IBM Germany(IBM德国分公司) ; FAST NUCES(FAST努赛斯大学) ; Askolay Pakistan(Askolay巴基斯坦)
专题命中 机器人数据与评测 :manipulation(title);分类 cs.CV、cs.LG
AI总结 TwoHead-SwinFPN通过双头架构和Swin Transformer结合FPN与Unet解码器,实现身份文档中合成操纵的高效检测与定位,具有高精度和计算效率。
Comments 8 pages