AI 中文总结
本研究提出轻量级策略层FABLE,通过在线学习用户执行策略实现LLM智能体个性化,在多类评估中改进偏好敏感行为且保持端到端任务性能。
AI 中文摘要
大型语言模型(LLM)智能体可检索记忆、调用工具、提出澄清问题并调整响应风格,但要将这些执行决策适配到单个用户仍存在困难。对专有系统而言,微调单独的LLM成本高昂或无法实现,而提示词和记忆主要是向智能体暴露用户信息,而非根据反馈调整其执行决策。我们将冻结智能体的个性化问题,表述为仅从已执行动作的标量反馈中学习每个用户执行策略的在线学习问题。我们提出FABLE(Factorized Adaptive Bandit Layer for Execution,执行用分解式自适应多臂老虎机层),这是一个位于潜在黑盒宿主智能体之外的轻量级策略层。FABLE将记忆、信息获取和响应决策进行分解,使反馈能更新相关选择;在探索前通过外部指定的可行集过滤动作;并通过贝叶斯上下文汤普森采样学习相对于固定默认-成本分数的用户特定残余偏好。在残余奖励线性模型下,校准后的变体继承了针对最佳可行动作的期望遗憾界。我们还刻画了在持续可行约束下无法识别的偏好,并提供了随时有效的虚假提升控制。在个性化推理、受控反馈和可执行工具使用评估中,FABLE相较于仅用规则的对照,改进了若干偏好敏感行为,同时在端到端任务性能上仍具有竞争力。
英文摘要
Large language model (LLM) agents can retrieve memory, call tools, ask clarifying questions, and vary response style, yet adapting these execution decisions to an individual user remains difficult. Fine-tuning a separate LLM is costly or impossible for proprietary systems, while prompts and memory primarily expose user information to the agent rather than adapt its execution decisions from feedback. We formulate personalization of a frozen agent as online learning of a per-user execution policy from scalar feedback observed only for the executed action. We propose FABLE (Factorized Adaptive Bandit Layer for Execution), a lightweight policy layer outside a potentially black-box host agent. FABLE factorizes memory, information-acquisition, and response decisions so feedback updates related choices; filters actions through an externally specified feasible set before exploration; and learns user-specific residual preferences relative to a fixed default-and-cost score via Bayesian contextual Thompson sampling. Under a linear residual-reward model, a calibrated variant inherits an expected-regret bound against the best feasible action. We also characterize preferences unidentifiable under persistent feasibility constraints and provide anytime-valid false-promotion control. Across personalized-reasoning, controlled-feedback, and executable tool-use evaluations, FABLE improves several preference-sensitive behaviors relative to rule-only control while remaining competitive on end-to-end task performance.