arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29548cs.LG

认证预测价值建议门控:强化学习中成本感知的语言模型指导

Certified Predictive Value-of-Advice Gating for Cost-Aware Language-Model Guidance in Reinforcement Learning

Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一种基于预测价值下界和动作证书的成本感知建议门控方法,在BabyAI上以显著减少的查询次数提升了回报,但逐状态优势未获充分验证。

中文摘要 AI 辅助

语言模型的建议可以加速强化学习,但调用成本高昂,且返回的动作可能过时或错误。我们将建议获取形式化为一个响应依存的元推理问题:在查询之前,控制器预测可能的解析响应,评估每个响应后可能采取的决策和声明的延续,并且仅当预测价值的下置信界超过定价成本时才进行查询。执行由单独的动作特定证书控制。在明确假设下,认证建议接近最优,包装学习器仅在干预稳定性下继承回退遗憾,且保守分配相对于短视预言机最多损失声明的查询价值估计误差。在BabyAI上,一个代理校准的控制器与Qwen2.5-1.5B和7B顾问相比,在不查询的情况下,GoToObj回报分别提高了0.029±0.016和0.030±0.015(20个种子),同时相对于始终查询,调用减少了97%以上。GoToLocal是一个无效结果。精确匹配调用的测试显示,仅1.5B顾问相对于随机放置有优势,而相对于等预算的早期调度没有优势。蒙德里安校准将决策相关的经验覆盖率从0.47提高到0.85,仍低于0.90的目标,而形式覆盖的半径是空洞的。因此,所展示的益处是在有用任务上的稳健稀疏建议量,而非经过验证的逐状态放置优势。

英文摘要

Language-model advice can accelerate reinforcement learning, but calls are costly and returned actions may be stale or wrong. We formulate advice acquisition as a response-contingent metareasoning problem: before querying, the controller predicts possible parsed responses, evaluates the decision and declared continuation that would follow each response, and queries only when a lower confidence bound on predictive value exceeds the priced cost. Execution is governed separately by an action-specific certificate. Under explicit assumptions, certified advice is near-optimal, a wrapped learner inherits fallback regret only under intervention stability, and conservative allocation loses at most the declared query-value estimation error relative to a myopic oracle. On BabyAI, a proxy-calibrated controller with Qwen2.5-1.5B and 7B advisors improves GoToObj return over no querying by 0.029 +/- 0.016 and 0.030 +/- 0.015 across 20 seeds while reducing calls by more than 97% relative to always-query. GoToLocal is a null result. Exactly matched-call tests show an advantage over random placement only for the 1.5B advisor and no advantage over an equal-budget early schedule. Mondrian calibration improves decision-relevant empirical coverage from 0.47 to 0.85, still below the 0.90 target, while the formally covered radius is vacuous. The demonstrated benefit is therefore robust sparse advice volume on a useful task, not a proven per-state placement advantage.

发表机构

  • Iowa State University(爱荷华州立大学)
  • Standard Chartered, Bangladesh(渣打银行(孟加拉国))
  • BRAC University(BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

↑