arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17247cs.AI

显式状态 elicitation 并不足够:对记忆-策略分类的受控审计

Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

Yihang Chen, Pin Qian, Su Wang, Chong Peng, Huan Xu, Shuaiting Li, Yiqi Sun

AI总结:

本文针对记忆-策略分类开展受控审计,发现显式状态 elicitation 不足以提升 Llama-3.3-70B 和 GPT-OSS-120B 的策略准确率,且样本级准确率夸大反事实一致性。

AI中文摘要:

个性化智能体必须决定是否使用、忽略、更新或查询检索到的用户记忆,以免其影响当前任务。我们利用该场景开发了针对结构化中间输出的实证审计协议:首先审计数据集 shortcut,然后隔离捆绑的提示变化,检查中间标签是否与答案相关,测试分解的语义证据,并审计提供者层面的执行失败。一个含480个样本的合成开发集最初显示,采用状态结构化提示 bundle 能带来较大提升,但 TF-IDF 诊断显示存在词汇可分性,且不存在独立的正向 Ignore 情况。因此我们构建了一个含160个样本的冻结受控反事实集,包含40个匹配的四元组家族及规则导出的参考策略。在该集合上,暴露四个状态定义可提升准确率,但孤立的显式状态-输出字段并未显著提升 Llama-3.3-70B 的策略准确率,仅为 GPT-OSS-120B 带来微弱且无统计显著性的提升。提供与基准相关的状态标签会改变策略预测,但由于这些标签确定性地映射到策略,这属于标签条件诊断而非忠实内部机制的证据。家族层面和种子稳定性分析进一步表明,样本级准确率夸大了反事实一致性:完整四元组家族成功的情况很少。一项探索性后续工作,即 elicitation 分解的语义证据,也未能提升干净评估端点的路由效果;对应的 GPT-OSS 条件因提供者侧请求验证不可用。我们仅评估策略分类,不评估下游响应、工具动作或记忆库突变。

英文摘要:

Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empirical audit protocol for structured intermediate outputs: first audit dataset shortcuts, then isolate bundled prompt changes, check whether intermediate labels are answer-associated, test decomposed semantic evidence, and audit provider-level execution failures. A 480-example synthetic development set initially suggested large gains from a state-structured prompt bundle, but TF-IDF diagnostics showed lexical separability and no positive standalone Ignore cases. We therefore construct a frozen 160-example controlled counterfactual set with 40 matched four-way families and rule-derived reference policies. On this set, exposing the four state definitions improves accuracy, but an isolated explicit state-output field does not significantly improve policy accuracy for Llama-3.3-70B and gives only a marginal, non-significant gain for GPT-OSS-120B. Supplying benchmark-associated state labels shifts policy predictions, but because those labels deterministically map to policies, this is a label-conditioning diagnostic rather than evidence of a faithful internal mechanism. Family-level and seed-stability analyses further show that example-level accuracy overstates counterfactual consistency: complete four-way family success is rare. An exploratory follow-up that elicits decomposed semantic evidence also fails to improve routing for the cleanly evaluated endpoint; the corresponding GPT-OSS condition was unavailable because of provider-side request validation. We evaluate policy classification only, not downstream responses, tool actions, or memory-store mutation.

补充信息

↑