TALK-Dem:痴呆相关沟通模式下的具身任务规划基准
TALK-Dem: Benchmarking Embodied Task Planning under Dementia-Associated Communication Patterns
浏览论文内容
中文总结 AI 辅助
提出TALK-Dem基准,评估LLM驱动机器人在痴呆相关沟通模式下的任务规划,并引入CARE方法提升鲁棒性,显著提高任务成功率。
中文摘要 AI 辅助
现有的基于大语言模型的机器人任务规划器依赖于一个理所当然的假设,即用户是理想的,其指令清晰、完整且以任务为中心。然而,在与现实世界用户(尤其是患有认知障碍的用户,如痴呆症患者)互动时,这些规划器常常出错,甚至带来身体安全风险。我们提出了TALK-Dem(痴呆中的言语属性和语言学知识),这是首个用于评估在痴呆相关言语沟通下基于大语言模型的机器人任务规划的基准。TALK-Dem包含4,800条指令,覆盖五种典型沟通模式,包括指称不精确、物体替代、空泛言语、话题漂移和侵入,每种模式有三个强度级别。在六个开放权重的大语言模型上的实验揭示了显著的鲁棒性差距。在各类沟通模式下,开放权重模型相比理想指令表现出高达22.3个百分点的性能下降。这揭示了现实应用中的关键差距甚至危险,尤其是在辅助机器人领域,由于隐私问题和连接限制,本地部署模型是必要的。为缓解这一问题,我们提出了上下文感知经验检索(CARE)方法,该方法检索先前已解决的相关任务,以提供任务特定的解释和规划上下文。CARE在六个开放权重模型上普遍优于标准提示基线,相比原始提示将平均任务成功率提高了18.1个百分点。这些结果强调了评估沟通鲁棒性和为本地部署的辅助机器人开发有效适应策略的重要性。TALK-Dem数据集可在以下网址公开获取:此https URL。
英文摘要
Existing LLM-driven robot task planners rely on a taken-for-granted assumption of an ideal user whose instructions are clear, complete, and task-focused. However, when interacting with real-world users, especially those experiencing cognitive impairments, such as people living with dementia (PLWD), the planners often make mistakes and even pose physical safety risks. We proposed TALK-Dem (Talking Attributes and Linguistic Knowledge in Dementia), the first benchmark for evaluating LLM-driven robot task planning under dementia-associated verbal communication. TALK-Dem contains 4,800 instructions and covers five typical communication patterns, including Referential Imprecision, Object Substitution, Empty Speech, Topic Drift, and Intrusion, at three intensity levels. Experiments across six open-weight LLMs reveal a substantial robustness gap. Across communication patterns, open-weight models exhibited performance drops of up to 22.3 percentage points compared to ideal instructions. This revealed a critical gap and even danger for real-world applications, especially in assistive robotics, where locally deployable models are necessary due to privacy concerns and connectivity constraints. To mitigate this issue, we proposed the Context-Aware Retrieval from Experience (CARE) method, which retrieves relevant previously resolved tasks to provide task-specific interpretation and planning context. CARE generally outperformed standard prompting baselines across the six open-weight models, improving average task success by 18.1 percentage points over the vanilla prompt. These results highlighted the importance of both evaluating communication robustness and developing effective adaptation strategies for locally deployable assistive robots. The TALK-Dem dataset is publicly available at https://anonymous.4open.science/r/TALK-Dem-A6B3/.
发表机构
- University of Notre Dame(圣母大学)
- Nanyang Technological University(南洋理工大学)
- Tohoku University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。