arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2025-12-24 至 2025-12-24 共收录 3 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 模仿学习与强化学习 3 篇

2507.04661 2025-12-24 cs.RO 85%

DRAE: Dynamic Retrieval-Augmented Expert Networks for Lifelong Learning and Task Adaptation in Robotics

DRAE:动态检索增强专家网络用于机器人终身学习与任务适应

Yayu Long, Kewei Chen, Long Jin, Mingsheng Shang

机构 * Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences(重庆绿色智能技术研究所,中国科学院)

专题命中 模仿学习与强化学习 :robotics(title,abstract);manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 DRAE通过动态检索增强专家网络实现机器人终身学习和任务适应,显著提升长期任务保留和知识重用能力。

Comments Accepted to the main conference of the Annual Meeting of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17180 2025-12-24 cs.RO cs.AI 74%

Conservative Bias in Multi-Teacher Learning: Why Agents Prefer Low-Reward Advisors

多教师学习中的保守偏见:为什么智能体更偏好低奖励顾问

Maher Mesto, Francisco Cruz

机构 * School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学院) Escuela de Ingeniería, Universidad Central de Chile(智利中央大学工程学院)

专题命中 模仿学习与强化学习 :navigation(abstract);robotic(abstract);分类 cs.RO、cs.AI;robotics(comments)

AI总结 本文揭示了多教师学习中智能体偏好保守低奖励顾问的现象,发现其受保守偏见主导,并在特定阈值下失效,同时在概念漂移情况下显著优于基线Q学习。

Comments 10 pages, 5 figures. Accepted at ACRA 2025 (Australasian Conference on Robotics and Automation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19717 2025-12-24 cs.LG cs.AI 62%

Thermodynamic Focusing for Inference-Time Search: Practical Methods for Target-Conditioned Sampling and Prompted Inference

推理时间搜索的热力学聚焦:目标条件采样和提示推理的实用方法

Zhan Zhang

机构 * Zhan Zhang(张湛)

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.AI、cs.LG

AI总结 本文提出ICFA算法,通过目标条件重新加权提升推理时间搜索效率,结合结构化提示与混合架构实现更高效的样本利用。

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏