arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27606cs.AI

状态锚定条件化:为方向依赖实时状态的面向用户LLM智能体提供包装

State-Grounded Conditioning: Wrapping User-Facing LLM Agents Where Direction Depends on Live State

Qi Liu, Xiaoyang Yuan, Yubin Ruan, Zhuomeng Zhang, Wenjin Wang, Di Wu, Mingye Xu, Xinyi Mou, Xingxi Yin, Ke Feng, Zixun Sun

首次发表
浏览论文内容

中文总结 AI 辅助

提出状态锚定条件化(SGC)设计原则,通过感知、锚定和交互包装器解决LLM智能体的方向漂移问题,在200会话基准上显著提升锚定准确率并降低延迟。

中文摘要 AI 辅助

我们提出状态锚定条件化(SGC),这是一种面向用户的LLM智能体的设计原则,此类智能体必须基于实时用户状态(游戏状态、会话历史、实时库存)进行条件化,同时我们识别出一类独特的失败模式,称为方向漂移:任务完成的响应所选择的方向与当前状态不一致。SGC通过感知、锚定和交互包装器,将状态依赖的控制外部化为针对结构化输入和三个主要状态切片的规则内核,并具有显式的条件化依赖关系。我们在一个包含200个会话的匿名基准(约1,000次助手模型轮次)上评估SGC,该基准来自一个游戏内对话教练智能体,用于指导玩家完成连续的竞技比赛,我们报告平均首令牌延迟和五项人工标注的对话质量指标,这些指标共同覆盖事实锚定和教练式指导进展。感知包装器将平均首令牌延迟保持在1.5秒(相比之下,生产工具使用框架内的PE-Agent为6.1秒);启用全部三个包装器将轮次级锚定准确率从61.1%/69.8%(提示/PE-Agent)提升至96.7%,会话级锚定准确率从20.0%/26.5%提升至83.5%;会话级锚定失败事件相对于最强基线减少约78%。累积消融实验显示,随着包装器的添加,产生了互补的增量收益。这些结果提示状态切片近似正交性,但未确立独立的各包装器效应。

英文摘要

We introduce State-Grounded Conditioning (SGC), a design principle for user-facing LLM agents that must condition on live user state (game state, session history, live inventory), and a distinct failure class we call direction drift: task-complete responses whose chosen direction misaligns with the current state. SGC externalises state-dependent control into rule kernels over structured inputs and three primary state slices, via Perception, Grounding, and Interaction wrappers with explicit conditioning dependencies. We evaluate SGC on a 200-session anonymised benchmark ($\approx$1,000 assistant model turns) from an in-game conversational coaching agent that guides players through consecutive competitive matches, reporting mean first-token latency and five human-annotated dialogue-quality metrics that jointly cover factual grounding and coach-like guidance progression. The Perception wrapper holds mean first-token latency at 1.5s (vs. 6.1s for PE-Agent inside a production tool-use harness); enabling all three wrappers lifts turn-level grounded accuracy from 61.1%/69.8% (Prompting / PE-Agent) to 96.7% and session-level grounded accuracy from 20.0%/26.5% to 83.5%; session-level grounding-failure incidents drop by $\approx$78% relative to the strongest baseline. A cumulative ablation shows complementary incremental gains as the wrappers are added. These results inform approximate state-slice orthogonality, without establishing independent per-wrapper effects.

发表机构

  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑