arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22035cs.RO

Ludi₀.₁:面向社交智能机器人的智能体系统

Ludi${}_{\scriptscriptstyle 0.1}$: An Agentic System for Socially Intelligent Robots

Wooseong Chung, William Cong, Jakub Dworakowski, Ethan Ewer, Tri Wahyu Guntara, Yeonwoo Jeong, Tianchong Jiang, Chaewon Kim, Hyunseo Kim, Jinwoo Kim, Jinyeon Ki… 展开作者

Wooseong Chung, William Cong, Jakub Dworakowski, Ethan Ewer, Tri Wahyu Guntara, Yeonwoo Jeong, Tianchong Jiang, Chaewon Kim, Hyunseo Kim, Jinwoo Kim, Jinyeon Kim, Yea-Seul Kim, Jack Kunde, Kangwook Lee, Sangheon Lee, Robert Nowak, Junha Roh

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出Ludi₀.₁智能体系统,集成多技能与微调视觉语言模型,为社交智能机器人提供可行的人机协作方案,还可生成相关交互轨迹用于开发更融合的机器人基础模型。

中文摘要 AI 辅助

机器人基础模型已在感知与控制领域取得显著进展,但自然的人机协作远不止执行孤立指令。机器人需识别歧义、维持多轮对话上下文、传达自身意图,并随用户意图变化调整当前行为。本文提出Ludi₀.₁,一款面向社交智能机器人的智能体系统,集成了交互式语音、多模态推理、记忆、导航及学习型操纵技能。其决策核心是经微调的视觉语言模型,该模型在涵盖歧义请求、澄清、修正、中断、混合社交与任务对话及多步骤任务的多轮交互轨迹上训练而成。专用管控模块负责管理模型与工具的交互循环,而专用导航与操纵策略则执行物理技能。Ludi₀.₁为实现流畅人机协作提供了可行路径,同时生成了开发人机深度融合机器人基础模型所需的多模态交互轨迹。

英文摘要

Robot foundation models have substantially advanced perception and control, but natural human-robot collaboration requires more than executing isolated commands. A robot must recognize ambiguity, maintain context across turns, communicate its intentions, and revise ongoing behavior as the user's intent changes. We present $\scriptstyle\mathsf{Ludi}_{\scriptscriptstyle 0.1}$, an agentic system for socially intelligent robots that integrates interactive speech, multimodal reasoning, memory, navigation, and learned manipulation. Its decision-making core is a fine-tuned vision-language model trained on multi-turn interaction traces spanning ambiguous requests, clarifications, corrections, interruptions, mixed social and task dialogue, and multi-step tasks. A purpose-built harness manages the model-tool interaction loop, while specialized navigation and manipulation policies execute physical skills. Ludi${}_{\scriptscriptstyle 0.1}$ demonstrates a practical path toward fluid human-robot collaboration today while producing the multimodal interaction traces needed to develop a more deeply integrated foundation model for robots and people.

发表机构

  • Ludo Robotics(乐动机器人)

机构由 AI 辅助整理,请以论文原文为准。

↑