AI 中文总结
本研究提出具身智能体中语言的五种功能角色,构建框架审查现有文献,发现语言功能使用与证据支持存在差距,以角色主张为单位可合理比较不同具身智能体。
AI 中文摘要
基础模型将语言融入具身智能体的各个环节,但语言的存在并不能表明其贡献了什么,也不能表明这种贡献的接地程度如何。本研究将这两个问题分开处理。我们定义了语言的五种非排他性功能角色:规范、具身表征、动作编排、接地调节以及执行耦合。对于每种角色,我们追踪从语言内容到其具身使用者的路径,并确定可用于检验所主张责任的观察或干预措施。将该框架应用于所综述的文献后,我们发现功能使用与证据支持之间存在反复出现的差距:可解释或经修订的语言中间产物可能不正确、未被使用,或无法影响后续行为;即使动作直接以语言为条件,系统层面的成功也无法单独分离出语言的贡献。因此,我们逐项评估接地主张,询问所报告的证据是否支持赋予语言的特定责任。以角色主张而非架构作为比较单位,使我们能够对模块化和端到端具身智能体进行比较,且不会将结论延伸至超出所报告证据的范围。
英文摘要
Foundation models place language throughout embodied agents, but its presence does not show what it contributes or how well that contribution is grounded. This survey separates these two questions. We define five non-exclusive functional roles for language: Specification, Embodied Representation, Action Orchestration, Grounding Regulation, and Execution Coupling. For each role, we trace the path from linguistic content to its embodied consumer and identify the observations or interventions that can test the claimed responsibility. Applying this framework to the reviewed literature reveals a recurring gap between functional use and evidential support. Interpretable or revised linguistic intermediates may be incorrect, go unused, or fail to affect later behavior. Even when actions are directly conditioned on language, system-level success does not by itself isolate language's contribution. We therefore evaluate grounding claim by claim, asking whether the reported evidence supports the specific responsibility assigned to language. Using role claims rather than architectures as the unit of comparison allows us to compare modular and end-to-end embodied agents without extending conclusions beyond the reported evidence.
Comments19 pages, 3 figures, 11 tables