arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2025-12-03 至 2025-12-03 共收录 46 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 机器人操作 14 篇

2512.02951 2025-12-03 cs.RO 83%

Experimental Characterization of Fingertip Trajectory following for a 3-DoF Series-Parallel Hybrid Robotic Finger

3自由度连杆驱动系列-并联混合机械手指尖轨迹跟踪的实验特性

Nicholas Baiata, Nilanjan Chakraborty

机构 * Department of Mechanical Engineering, Stony Brook University(机械工程系,石英布鲁克大学)

专题命中 机器人操作 :robotic(title,abstract);manipulation(abstract);分类 cs.RO

AI总结 本文提出了一种三自由度连杆驱动的机器人手指,通过实验验证了其在任务空间轨迹跟踪中的高精度性能,为灵巧手部操作提供了新的基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06183 2025-12-03 cs.RO 83%

Sampling-Based Model Predictive Control for Dexterous Manipulation on a Biomimetic Tendon-Driven Hand

基于采样的模型预测控制在仿生腱驱动手的灵巧操作中的应用

Adrian Hess, Alexander M. Kübler, Benedek Forrai, Mehmet Dogar, Robert K. Katzschmann

机构 * Soft Robotics Lab, Department of Mechanical and Process Engineering, ETH Zurich(苏黎世联邦理工学院机械与过程工程系软机器人实验室) School of Computer Science, University of Leeds(利兹大学计算机科学学院)

专题命中 机器人操作 :manipulation(title,abstract);robotic(abstract);分类 cs.RO

AI总结 本研究提出利用基于采样的模型预测控制实现仿生腱驱动手的灵巧操作,通过结合视觉语言模型与实时优化器,有效生成接触丰富的行为。

Comments For a video, see https://youtu.be/u4d6v3ohsOI

Journal ref 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01629 2025-12-03 cs.CV cs.RO 82%

SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge

SPARK: 面向模拟的关节化重建与VLK知识

Yumeng He, Ying Jiang, Jiayin Lu, Yin Yang, Chenfanfu Jiang

专题命中 机器人操作 :robotics(abstract);embodied AI(abstract);manipulation(abstract);robotic(abstract)

AI总结 SPARK通过结合VLK和生成扩散模型,实现从单张图像中生成物理一致的关节化3D物体,提升机器人操作和交互建模的应用效果。

Comments Project page: https://heyumeng.com/SPARK/index.html. 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05858 2025-12-03 cs.RO 79%

ViTaMIn-B: A Reliable and Efficient Visuo-Tactile Bimanual Manipulation Interface

ViTaMIn-B: 一种可靠且高效的视觉-触觉双臂操作接口

Chuanyu Li, Chaoyi Liu, Daotan Wang, Shuyu Zhang, Lusong Li, Zecui Zeng, Fangchen Liu, Jing Xu, Rui Chen

机构 * Tsinghua University(清华大学) University of California, Berkeley(加州大学伯克利分校) JD Explore Academy(JD探索学院) The Hong Kong Polytechnic University(香港理工大学)

专题命中 机器人操作 :manipulation(title,abstract);分类 cs.RO

AI总结 ViTaMIn-B通过DuoTact传感器和6自由度双臂位姿采集技术,实现了高效可靠的双臂操作任务数据采集。

Comments Project page: https://chuanyune.github.io/ViTaMIn-B_page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02609 2025-12-03 cs.RO cs.CV 62%

SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction

SAM2Grasp:通过提示条件化的时序动作预测解决多模态抓取

Shengkai Wu, Jinrong Yang, Wenqiu Luo, Linfeng Gao, Chaohui Shang, Meiyu Zhi, Mingshan Sun, Fangping Yang, Liangliang Ren, Yong Zhao

机构 * CVTE(中国中车)

专题命中 机器人操作 :robotic(abstract);分类 cs.RO、cs.CV

AI总结 SAM2Grasp通过提示条件化的时序动作预测,解决多模态抓取中的冲突问题,实现高精度抓取性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21025 2025-12-03 cs.SD cs.AI cs.LG eess.AS 62%

Text-Queried Audio Source Separation via Hierarchical Modeling

通过分层建模实现文本查询的音频源分离

Xinlei Yin, Xiulian Peng, Xue Jiang, Zhiwei Xiong, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) School of Information and Communication Engineering, Communication University of China(中国通信大学信息与通信工程学院) Microsoft Research Asia(微软亚洲研究院)

专题命中 机器人操作 :manipulation(abstract);分类 cs.AI、cs.LG

AI总结 本文提出HSM-TSS框架,通过分层建模实现文本查询的音频源分离,结合双阶段语义分离与结构保持重建,提升复杂场景下的分离性能与语义一致性。

Comments Accepted by TASLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02422 2025-12-03 quant-ph cs.AI cs.LG 62%

Quantum feature encoding optimization

量子特征编码优化

Tommaso Fioravanti, Brian Quanz, Gabriele Agliardi, Edgar Andres Ruiz Guzman, Ginés Carrascal, Jae-Eun Park

机构 * IBM Quantum(IBM量子实验室) IBM Italy(IBM意大利)

专题命中 机器人操作 :manipulation(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过优化数据编码方式提升量子机器学习模型性能,通过经典数据处理预处理步骤并结合真实量子硬件实验验证有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01856 2025-12-03 cs.RO 57%

Is Image-based Object Pose Estimation Ready to Support Grasping?

基于图像的对象姿态估计是否准备好支持抓取?

Eric C. Joyce, Qianwen Zhao, Nathaniel Burgdorfer, Long Wang, Philippos Mordohai

机构 * Stevens Institute of Technology(史蒂文斯理工学院)

专题命中 机器人操作 :robotic(abstract);分类 cs.RO

AI总结 本文提出评估基于图像的对象姿态估计器的框架,通过模拟实验验证其在机器人抓取中的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17498 2025-12-03 cs.AI cs.CL cs.NE cs.SC 57%

Mechanisms of Symbol Processing for In-Context Learning in Transformer Networks

Transformer网络中上下文学习符号处理的机制

Paul Smolensky, Roland Fernandez, Zhenghao Herbert Zhou, Mattia Opper, Adam Davies, Jianfeng Gao

专题命中 机器人操作 :manipulation(abstract);分类 cs.AI

AI总结 本文提出了一种基于生产系统语言PSL的转换器架构,用于提升转换器在符号处理中的能力,展示了其在抽象符号任务中的应用和图灵通用性。

Journal ref Journal of Artificial Intelligence Research, 84(23) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01783 2025-12-03 cs.LG cs.GT 57%

The Active and Noise-Tolerant Strategic Perceptron

主动且抗噪声的战略感知机

Maria-Florina Balcan, Hedyeh Beyhaghi

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 机器人操作 :manipulation(abstract);分类 cs.LG

AI总结 本文提出了一种在战略环境下主动学习线性分离器的算法,通过减少标签查询数量实现高效分类,即使在数据不一致的情况下也能保持性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02961 2025-12-03 physics.optics 50%

Reflective Metalenses for Near-Infrared Wavelengths Based on Silicon Nanorods

基于硅纳米棒的近红外波长反射型金属透镜

Iftekhar Ahmed, K. B. M. Sharif Mahmood, Tanvin Tamanna

专题命中 机器人操作 :manipulation(abstract)

AI总结 本文提出了一种基于硅纳米棒的近红外反射型金属透镜设计,通过几何优化实现2π相位控制,用于提升光聚焦和操控性能,适用于成像、通信和传感领域。

Comments 4 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02658 2025-12-03 physics.optics 50%

Multimode interface between optical free-space- and waveguide modes

光自由空间与波导模式之间的多模式接口

Teresia Stranden, Oussama Korichi, Matias Eriksson, Matteo Cherchi, George Thomas, Robert Fickler

专题命中 机器人操作 :manipulation(abstract)

AI总结 本文提出一种多平面光转换方案,实现自由空间高阶LG模式与波导模式的高效宽频带转换,为多模式光通信和芯片处理提供新途径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02263 2025-12-03 cs.HC cs.GR 50%

DepthScape: Authoring 2.5D Designs via Depth Estimation, Semantic Understanding, and Geometry Extraction

DepthScape: 通过深度估计、语义理解和几何提取进行2.5D设计创作

Xia Su, Cuong Nguyen, Matheus A. Gadelha, Jon E. Froehlich

专题命中 机器人操作 :manipulation(abstract)

AI总结 DepthScape通过深度估计、语义理解和几何提取,实现2.5D设计创作,提升视觉真实感与动态效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15519 2025-12-03 cond-mat.mes-hall physics.app-ph 50%

Correlating on-the-fly Electrical and Optical Skyrmion Readout

实时电与光学Skyrmion读出相关性

Grischa Beneke, Kilian Leutner, Nikhil Vijayan, Fabian Kammerbauer, Duc Minh Tran, Sachin Krishnia, Johannes Güttinger, Armin Satz, Robert Frömter, Mathias Kläui

专题命中 机器人操作 :manipulation(abstract)

AI总结 本研究提出一种利用热激活Skyrmions的实时电与光学读出方法,通过霍尔电压与Kerr显微镜成像的实时相关性,实现对单个移动Skyrmions的可靠检测,并推导出适用于微米到纳米尺度的分析公式。

Comments 11 pages, 4 figures

Journal ref Appl. Phys. Lett. 127, 222403 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身导航 11 篇

2504.09843 2025-12-03 cs.CV cs.RO 81%

ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments

ST-Booster: 一种用于连续环境视觉-语言导航的迭代时空感知增强器

Lu Yue, Dongliang Zhou, Liang Xie, Erwei Yin, Feitian Zhang

机构 * Robotics and Control Laboratory, the School of Advanced Manufacturing and Robotics, and the State Key Laboratory of Turbulence and Complex Systems, Peking University(机器人与控制实验室、先进制造与机器人学院、湍流与复杂系统国家重点实验室,北京大学) Defense Innovation Institute, Academy of Military Sciences(国防创新研究院,军事科学院) Tianjin Artificial Intelligence Innovation Center(天津人工智能创新中心) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 ST-Booster通过多粒度感知和指令感知推理提升连续环境中视觉-语言导航的性能。

Comments 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11590 2025-12-03 cs.RO cs.LG 81%

Predicting Human Perceptions of Robot Performance During Navigation Tasks

预测人类在导航任务中对机器人性能的感知

Qiping Zhang, Nathan Tsoi, Mofeed Nagib, Booyeon Choi, Jie Tan, Hao-Tien Lewis Chiang, Marynel Vázquez

机构 * Yale University(耶鲁大学) Google DeepMind, Google Inc(谷歌DeepMind及谷歌公司)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.LG

AI总结 本文通过SEAN TOGETHER数据集研究了利用非语言行为线索和机器学习预测人类对机器人导航性能感知的方法,发现空间特征推理和监督学习在预测中表现更优。

Journal ref ACM Transactions on Human-Robot Interaction, Vol. 14, No. 3, Article 46, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02423 2025-12-03 cs.CV 79%

GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning

GUI探索实验室:通过多轮强化学习提升代理的屏幕导航

Haolong Yan, Yeqing Shen, Xin Huang, Jia Wang, Kaijun Tan, Zhixuan Liang, Hongxin Li, Zheng Ge, Osamu Yoshie, Si Li, Xiangyu Zhang, Daxin Jiang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) StepFun Waseda University(早稻田大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 具身导航 :navigation(title,abstract);分类 cs.CV

AI总结 GUI探索实验室通过多轮强化学习提升代理屏幕导航能力,结合监督微调和交互式试错,实现更高效和通用的GUI代理训练。

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02400 2025-12-03 cs.CV 79%

Nav-$R^2$ Dual-Relation Reasoning for Generalizable Open-Vocabulary Object-Goal Navigation

Nav-$R^2$双关系推理用于通用开放词汇目标导航

Wentao Xiang, Haokang Zhang, Tianhang Yang, Zedong Chu, Ruihang Chu, Shichao Xie, Yujian Yuan, Jian Sun, Zhining Gu, Junjie Wang, Xiaolong Wu, Mu Xu, Yujiu Yang

机构 * Tsinghua University(清华大学) Amap, Alibaba Group(阿里巴巴集团的阿里的地图)

专题命中 具身导航 :navigation(title,abstract);分类 cs.CV

AI总结 Nav-R2通过双关系推理和相似性感知记忆,在开放词汇环境下实现高效目标导航,提升未见对象定位性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16904 2025-12-03 cs.RO cs.AI 73%

Deep Learning for Human Locomotion Analysis in Lower-Limb Exoskeletons: A Comparative Study

深度学习在下肢外骨骼中的人体运动分析:比较研究

Omar Coser, Christian Tamantini, Matteo Tortora, Leonardo Furia, Rosa Sicilia, Loredana Zollo, Paolo Soda

专题命中 具身导航 :robotics(abstract);robotic(abstract);分类 cs.RO、cs.AI

AI总结 本文通过比较八种深度神经网络架构,研究了在不同地形中预测人体运动参数的方法,并展示了使用IMU传感器的高效和轻量级系统设计。

Comments 26 pages, 6 figures

Journal ref Front. Comput. Sci., 7:1597143 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21214 2025-12-03 cs.CV cs.CL 70%

VoxRep: Enhancing 3D Spatial Understanding in 2D Vision-Language Models via Voxel Representation

VoxRep:通过体素表示增强2D视觉-语言模型的3D空间理解

Alan Dao, Norapat Buppodom

机构 * Menlo Research(Menlo研究)

专题命中 具身导航 :robotics(abstract);navigation(abstract);分类 cs.CV

AI总结 本文提出VoxRep方法,通过将体素空间切分为2D切片并输入预训练的视觉-语言模型,实现对3D环境的高效语义理解。

Journal ref Proc. APSIPA ASC 2025, pp. 1464-1469

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02389 2025-12-03 cs.AI cs.LG 62%

Synthetic Error Injection Fails to Elicit Self-Correction In Language Models

合成错误注入无法在语言模型中引发自我修正

David X. Wu, Shreyas Kapur, Anant Sahai, Stuart Russell

机构 * ucb(伯克利大学)

专题命中 具身导航 :robotics(abstract);分类 cs.AI、cs.LG

AI总结 本研究发现合成错误注入无法有效提升语言模型的自我修正能力,揭示了策略性强化学习在激发自我修正方面的独特优势。

Comments 13 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02994 2025-12-03 eess.SY cs.SY eess.SP 56%

GNSS Array-Based Multipath Detection Employing UKF on Manifolds

基于GNSS阵列的UKF在流形上实现多路径检测

Abdelgabar Ahmed, Tarig Ballal, Xing Liu, Mohanad Ahmed, Tareq Y. Al-Naffouri

专题命中 具身导航 :navigation(abstract,comments)

AI总结 本文提出基于GNSS阵列和UKF在流形上的多路径检测方法,利用RANSAC算法提高计算效率,有效提升定位精度。

Comments The paper, was presented at the ION PLANS 2025 meeting (Position, Location, and Navigation Symposium) in Session C1: Multisensor Integrated Systems and Sensor Fusion Technologies, and is published in the conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15286 2025-12-03 math.NA cs.NA 50%

Real-time aerodynamic load estimation for hypersonics via strain-based inverse maps

通过应变基逆映射实现高超声速的实时气动载荷估计

Julie Pham, Omar Ghattas, Noel Clemens, Karen Willcox

专题命中 具身导航 :navigation(abstract)

AI总结 本文提出通过应变基逆映射实时估计高超音速飞行器气动载荷的方法,适用于恶劣环境下的压力推断与力矩系数计算。

Comments 22 pages, 15 figures

Journal ref AIAA Journal, Vol. 63, No. 1, 2025, pp. 91-101

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12766 2025-12-03 cs.CR 50%

GPS-Spoofing Attack Detection Mechanism for UAV Swarms

无人机群中GPS欺骗攻击检测机制

Pavlo Mykytyn, Marcin Brzozowski, Zoya Dyka, Peter Langendoerfer

专题命中 具身导航 :navigation(abstract)

AI总结 本研究提出了一种基于UAV群成员间距离对比的GPS欺骗攻击检测机制,通过脉冲无线电超宽带测距技术识别单发射机和多发射机攻击,以防止UAV被误导或坠毁。

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02043 2025-12-03 cs.CL 50%

Mirror, Mirror on the Wall -- Which is the Best Model of Them All?

镜中镜——哪一个模型才是最好的?

Dina Sayed, Heiko Schuldt

机构 * Databases and Information Systems Research Group University of Basel, Switzerland(数据库与信息系统研究组巴塞尔大学)

专题命中 具身导航 :navigation(abstract)

AI总结 本文提出了一种模型选择方法论,通过分析排行榜和基准来帮助选择最适合特定用途的模型。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 具身推理 4 篇

2511.19584 2025-12-03 cs.LG cs.CV cs.RO 82%

Learning Massively Multitask World Models for Continuous Control

学习大规模多任务世界模型用于连续控制

Nicklas Hansen, Hao Su, Xiaolong Wang

机构 * University of California San Diego(加州大学圣地亚哥分校)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.CV、cs.LG

AI总结 Newt通过大规模预训练和在线交互,实现智能体在数百个任务上的多任务学习,提升连续控制的效率和适应性。

Comments Webpage: https://www.nicklashansen.com/NewtWM

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02417 2025-12-03 cs.RO cs.AI 81%

Vehicle Dynamics Embedded World Models for Autonomous Driving

用于自动驾驶的车辆动力学嵌入世界模型

Huiqian Li, Wei Pan, Haodong Zhang, Jin Huang, Zhihua Zhong

机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院) Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系) Tsinghua University, Chinese Academy of Engineering(清华大学、中国工程院)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出VDD方法,通过分离车辆动力学与环境动力学建模,提升自动驾驶中对车辆参数变化的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02793 2025-12-03 cs.CV 79%

IC-World: In-Context Generation for Shared World Modeling

IC-World:基于上下文的共享世界建模生成

Fan Wu, Jiacheng Wei, Ruibo Li, Yi Xu, Junyou Li, Deheng Ye, Guosheng Lin

机构 * Nanyang Technological University(南洋理工大学) Goertek Alpha Labs(歌尔声学实验室) Tencent(腾讯)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 IC-World通过激活大视频模型的上下文生成能力,实现了多视角共享世界建模,通过强化学习和奖励模型提升生成视频的几何和运动一致性。

Comments codes:https://github.com/wufan-cse/IC-World

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02419 2025-12-03 q-bio.NC cs.AI cs.CL cs.NE 79%

The brain-AI convergence: Predictive and generative world models for general-purpose computation

脑与人工智能的融合:面向通用计算的预测性和生成性世界模型

Shogo Ohmae, Keiko Ohmae

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本文探讨脑与AI在世界模型计算中的共同机制,揭示预测性与生成性模型如何实现通用计算及类人智能。

Comments 22 pages, 4 figures. Related to our earlier preprint "The brain versus AI" (arXiv:2411.16075) but a distinct article. The earlier work surveyed broad brain-AI parallels; here we focus on world-model-based computation and convergent evolution between the brain and AI, especially large language models

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 模仿学习与强化学习 3 篇

2512.02022 2025-12-03 cs.RO 83%

Reinforcement Learning for Robotic Safe Control with Force Sensing

基于力感知的机器人安全控制强化学习

Nan Lin, Linrui Zhang, Yuxuan Chen, Zhenrui Chen, Yujun Zhu, Ruoxi Chen, Peichen Wu, Xiaoping Chen

专题命中 模仿学习与强化学习 :robotic(title,abstract);manipulation(abstract);分类 cs.RO

AI总结 本文提出基于力感知的强化学习方法,提升机器人在复杂环境中的安全性和可靠性,通过实验验证其在物体推动任务中的高效性与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏