arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24888cs.AI

SIMGUIDE:用于个性化智能体规划的过程式基础多上下文表示

SimGuide: Typed Multi-Context User Representations for Preference-Conditioned Agent Planning

Chirag Shah

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对个性化AI智能体无法适配用户多场景优先级的问题,提出SIMGUIDE方法,构建SIMBENCH基准验证,发现过程式基础的Sims优于RAG,且表示格式是关键设计变量。

中文摘要 AI 辅助

个性化AI智能体大多将用户视为单一实体,即拼接为提示词的扁平画像。当同一用户在不同生活场景中持有不同优先级时,这种方式会失效;而当这些优先级冲突时,失效会极其严重。核心问题并非智能体缺乏用户信息,而是用户表示的格式决定了智能体能否利用这些信息行动。我们提出SIMGUIDE,一种将用户上下文结构化称为Sims的类型化、领域特定模块的方法,并通过过往决策的过程示例为每个约束提供依据。为评估该方法,我们构建了SIMBENCH,这是一个包含47项偏好条件规划任务的诊断套件,其中正确规划取决于激活的用户上下文——这是现有基准均未测试的特性。仅声明式Sim约束的表现未优于基于检索的个性化(RAG)。过程式基础的Sims在GPT-4o上的表现优于RAG(偏好 adherence 提升7.9个百分点,p=0.013),且在GPT-4o和Claude Sonnet 4.5的100项τ-bench任务中均复现了该优势(p≤0.023)。在参数层面,相同原则成立:训练分布决定参数适配能否成功。任务匹配的LoRA微调相比未适配的基础模型,将生成质量提升12.8个ROUGE-L点,而按Sim类型而非用户身份路由适配器,进一步提升7.3个点,且对28%的路由错误具有鲁棒性。表示格式——而非表示内容——是首要设计变量。

英文摘要

Agents that act on a user's behalf must plan differently for different users, and increasingly do so from some structured representation of user context and not from raw interaction history. How much that structure is worth, and which parts of it carry the value, is largely unmeasured. We introduce SimBench, 47 preference-conditioned planning tasks over 9 synthetic users represented as 28 typed, potentially conflicting context blocks, where the correct plan depends on which contexts are active and how their conflicts are resolved. Against it we evaluate SimGuide, a framework combining typed multi-context representation, explicit conflict arbitration, and optional procedural grounding of individual constraints. Across three models and six user-context representations, SimGuide's typed blocks with arbitration outperform retrieval over the same user's past decisions by +0.210, +0.205 and +0.144 Preference Adherence on Llama 3.3 70B, GPT-4o and Claude Sonnet 4.5 respectively (all p < 0.001); removing the arbitration instruction alone costs up to +0.209. Grounding each constraint with a worked example of past application helps only where the model has headroom: +0.094 on Llama 70B (p < 0.001), falling to +0.029 on GPT-4o and +0.002 on Claude, which already scores perfectly on 35 of 47 tasks without it. We report the benchmark's minimum detectable effect alongside its results. The benchmark ships with a provenance audit that re-derives every reported number from the prompt that produced it.

发表机构

  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑