发表机构
Stony Brook University; Meta Platforms(石溪大学; Meta平台)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语言反馈易被误解导致排除有效解的问题,提出TRACE算法,将反馈学习建模为约束树上的纯探索,通过证伪与识别区分反馈用法,显著提升任务成功率与鲁棒性。
AI 中文摘要
在交互式学习中,自然语言反馈通常通过指出违反的要求来解释动作失败的原因。误解这种反馈可能导致智能体排除有效解决方案。我们通过将用户意图建模为动作空间上的潜在约束,并将从语言反馈中学习表述为对可行区域的纯探索来研究这一设置。我们提出了TRACE算法,该算法将候选约束组织成树结构,并通过生成满足该约束的动作来测试每个提议的细化。TRACE仅在重复测试中反馈不与其矛盾时才提交该细化。我们区分了使用相同反馈的两种方式:(i)证伪,检测与当前测试约束集的矛盾;(ii)识别,可能额外指出违反的约束。我们证明了TRACE-Falsification的高概率覆盖界,其依赖于候选类大小$H$。在可靠识别下,TRACE-Identification可以将此依赖替换为$K/p_{\mathrm{ext}}$,其中$K$是潜在约束的数量,$p_{\mathrm{ext}}$是从信息性反馈中提取缺失真实约束的概率下界。我们在六个语言反馈任务上评估了TRACE。在RecMovie上,TRACE-Identification在评估输出上限为20和60时分别达到73%和86%的最终输出成功率,而在相同反馈和输出上限下,评估的提示基线最多仅达到42%和48%。受控的身份损坏实验进一步表明,当证伪检测器保持可靠时,TRACE比直接累积具有更强的鲁棒性。
英文摘要
Natural-language feedback in interactive learning often explains why an action failed by pointing to violated requirements. Misinterpreting this feedback can lead an agent to rule out valid solutions. We study this setting by modeling user intent as latent constraints over an action space and formulating learning from language feedback as pure exploration over feasible regions. We introduce TRACE, an algorithm that organizes candidate constraints in a tree and tests each proposed refinement by generating actions that satisfy it. TRACE commits to the refinement only if the resulting feedback does not contradict it over repeated tests. We distinguish two ways of using the same feedback: (i) falsification, which detects contradictions to the constraint set currently being tested, and (ii) identification, which may additionally name a violated constraint. We prove high-probability coverage bounds with dependence on the candidate class size $H$ for TRACE-Falsification. With reliable identification, TRACE-Identification can replace this dependence by $K/p_{\mathrm{ext}}$, where $K$ is the number of latent constraints and $p_{\mathrm{ext}}$ lower-bounds the probability of extracting a missing true constraint from informative feedback. We evaluate TRACE across six language-feedback tasks. On RecMovie, TRACE-Identification achieves 73% and 86% final-output success under caps of 20 and 60 evaluated outputs, compared with at most 42% and 48% for the evaluated prompting baselines given the same feedback and output caps. Controlled identity-corruption experiments further show greater robustness than direct accumulation when the falsification detector remains reliable.