发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出AALT方法,通过构建潜在枢纽拓扑并请求最大化起始-目标连通性增益的桥梁演示,在72任务模拟中仅用3个演示即达100%成功率。
AI 中文摘要
主动模仿学习通过允许学习器请求它需要的演示来减少专家的工作量。现有方法通常根据这些请求对专家策略的预期信息增益来选择。然而,在结构化多任务领域中,尽管任务的解决方案共享可复用的行为,但起始-目标任务的数量可能组合式增长。这使得可组合行为特别有价值,因为单个演示可能有助于同时解决许多任务。先前的方法在选择请求哪个演示时没有明确考虑这一价值。我们引入了通过潜在拓扑的自适应智能体(AALT),它请求那些能最大化起始-目标连通性预期增益的演示。我们进一步表明,该目标在形式上与关于任务可达性的信息增益相关。AALT将现有演示组织成一个由学习行为连接的潜在枢纽状态拓扑,识别出可能同时启用许多任务的高价值桥梁演示,并将每个演示落实到专家查询中。在推理时,它通过生成的拓扑进行规划,并让扩散策略以每个连续的枢纽转换为条件。在一个模拟的UR5e机器人有序检索领域(包含72个任务)中,AALT仅使用3个演示(总共5个转换,超出初始数据集)就持续地将成功任务从42/72提高到72/72(100%)。在20个演示后,最强的基线平均使用98个转换达到88.6%的成功率。
英文摘要
Active imitation learning reduces expert effort by allowing a learner to request the demonstrations it needs. Existing methods typically select these requests for their expected information gain about the expert policy. In structured multi-task domains, however, the number of start-goal tasks may grow combinatorially despite their solutions sharing reusable behavior. This makes composable behaviors especially valuable, since a single demonstration may help solve many tasks at once. Prior methods do not explicitly account for this value when selecting which demonstration to request. We introduce Adaptive Agents via Latent Topologies (AALT), which requests demonstrations that maximize expected gains in start-goal connectivity. We further show that this objective is formally tied to information gain about task reachability. AALT organizes existing demonstrations into a topology of latent hub states connected by learned behaviors, identifies high-value bridge demonstrations that are likely to enable many tasks at once, and grounds each to an expert query. At inference, it plans through the resulting topology and conditions a diffusion policy on each successive hub transition. In a simulated UR5e robot ordered-retrieval domain with 72 tasks, AALT improved from 42/72 to 72/72 (100%) successful tasks consistently using only 3 demonstrations totaling 5 transitions beyond the initial dataset. After 20 demonstrations, the strongest baseline averaged 88.6% success using 98 transitions.