arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

库存盘点:使用公平预言机测量大语言模型智能体中感知与行动之间的差距

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

Sagar Deb, Ashwanth Krishnan

arXiv 2607.13618首次发表:更新:

发表机构

QpiAI(QpiAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型智能体在多周决策任务中的知行差距,提出库存盘点基准,通过公平参考策略区分感知与行动失败,测量知行差距两方向,为评估大语言模型智能体提供新途径。

AI 中文摘要

大语言模型智能体越来越多地在多周决策任务中接受评估,在这些任务中,驱动成本的状态从未被直接观察到。在这样的任务中,最终成本无法说明智能体失败的原因:它可能误解了世界,或者正确地理解了世界但仍然未能采取行动(知行差距)。现有评估无法区分这两种失败;它们的参考策略要么读取智能体从未见过的特权信息,要么完全缺失。我们引入了库存盘点(STOCKTAKE),这是一个为期26周的供应链补货基准,构建为具有六个隐藏因素过程的因子部分可观测马尔可夫决策过程,设计使得公平的参考策略是可计算的:每个因素的精确贝叶斯滤波器驱动在智能体接收到的相同观测流上的展开策略。在症状盲基本库存下限(0)和这个预言机(1)之间对每次运行进行评分会产生一个技能分数,对每周的书面理由进行评分会产生一个陈述信念检测滞后和一个知行率,从而分别测量状态估计和控制。在五十个具有精心策划的压力配置文件的种子上,Claude Sonnet 5、GPT - 5.4、DeepSeek - V4 - Pro和Grok 4.将库存盘点(STOCKTAKE)应用于多个大语言模型智能体在供应链补货任务的评估中,通过公平可计算的参考策略区分感知与行动失败,测量知行差距的两个方向,为评估大语言模型智能体提供新方法。

英文摘要

LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read it correctly and still failed to act (the knowing-doing gap). Existing evaluations cannot separate these two failures; their reference policies either read privileged information the agent never sees, or are missing altogether. We introduce STOCKTAKE, a 26-week supply-chain replenishment benchmark built as a factored partially observable Markov decision process with six hidden factor processes, designed so that a fair reference policy is computable: an exact Bayes filter per factor drives a rollout policy on the identical observation stream the agent receives. Scoring each run between a symptom-blind base-stock floor (0) and this oracle (1) yields a skill score, and grading each week's written rationale yields a stated-belief detection lag and a knowing-doing rate, so state estimation and control are measured separately. On fifty seeds with curated stress profiles, Claude Sonnet 5, GPT-5.4, DeepSeek-V4-Pro, and Grok 4.5 detect 84-88% of hidden failures, typically within a week of onset, yet span skill scores from 0.62 to -0.23: two of the four end below the symptom-blind floor while naming factors slightly faster than the two that beat it. The failure has two faces. Where stress persists, 34-43% of correctly diagnosed stress weeks still end in stockout for every model, a rate that partly reflects the severity of the weeks models notice. That rate also runs opposite to skill: the two models under the floor stock out least on diagnosed weeks, so under-response is only one face of the gap, and their traces point to the other, responses whose cost exceeds what they protect. STOCKTAKE measures both directions of that failure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑