arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重放差距:大语言模型智能体中模型切换的静态评估评估的是错误的世界

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World

Ashritha Gonuguntla

arXiv 2608.08239首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现大语言模型智能体的模型切换静态评估基于错误假设,通过分支展开实验证明模型交换会导致轨迹分歧,重放评估无法准确预测结果,发布了相关工具与轨迹。

AI 中文摘要

大语言模型(LLM)路由器通过将每个请求匹配到成本最低的合适模型来保证效率,且正越来越多地应用于多步智能体的每一步中。然而,智能体路由器的评估方式与单轮路由器相同:通过重放记录的轨迹并替换另一个模型的记录输出,假设轨迹的其余部分不受影响。我们通过分支展开测试这一假设:在受控点分叉实时 SWE-bench 智能体轨迹,重建环境,用不同模型继续每个分叉,并与相同模型的对照分叉进行比较,以分离采样和重放噪声。在六次配对运行(约900次展开)中,交换操作超出其匹配对照基线的标准化编辑距离为+0.25至+0.66(多重性校正置信区间排除零),重写了分叉后61-94%的动作;74-77%的早期交换在分叉后第一个动作就发生分歧,而对照分叉仅为6-35%,仅3%的重放状态有效。两种情况下分歧均随分叉深度增加而降低。我们观察到的所有5次结果翻转均出现在交换分支中,升级操作挽救了未解决的实例,降级操作则丢失了唯一的解决方案,而359次对照分叉中未出现任何结果翻转。使用日志拼接重放评估器对这些相同交换进行评分时,重放错误预测了所有与成功相关的结果调用,且预测补丁与实际情况的相似度为0.00-0.11。对噪声基线的审计显示,温度为0的“确定性”依赖于配置:FP8服务的对照分叉在超过90%的分叉上发生分歧,而AWQ服务的对照分叉则保持几乎相同;在严格预算下,更强的模型更常耗尽其步骤而不提交。基于重放的基准对智能体路由评估的是错误的世界;我们发布了我们的工具和所有轨迹。

英文摘要

LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged trajectories and substituting another model's recorded outputs, assuming the rest of the trajectory is unaffected. We test this assumption with branching rollouts: we fork live SWE-bench agent trajectories at controlled points, rebuild the environment, continue each fork with a different model, and compare against same-model control forks that isolate sampling and replay noise. Across six paired runs (~900 rollouts), swaps exceed their matched control floors by +0.25 to +0.66 normalized edit distance (multiplicity-corrected CIs exclude zero), rewriting 61-94% of post-fork actions; 74-77% of early swaps diverge at the first post-fork action, versus 6-35% of controls, leaving only 3% of replayed states valid. Divergence decreases with fork depth in both directions. All five outcome flips we observe occur in swap arms, upgrades rescuing unsolved instances and a downgrade losing the sole solve, and zero occur across 359 control forks. Scoring these same swaps with a log-stitching replay evaluator, replay mispredicts every success-relevant outcome call and predicts patches with 0.00-0.11 similarity to reality. Auditing the noise floor, temperature-0 "determinism" is configuration-dependent: FP8-served controls diverge on over 90% of forks while AWQ-served ones remain near-identical; and under tight budgets the stronger model more often exhausts its steps without submitting. Replay-based benchmarks score the wrong world for agentic routing; we release our harness and all trajectories.

Comments8 pages, 3 figures. Accepted at the Conference on Language Modeling 2026. Code: https://github.com/AshrithaG/replay-gap Data: https://huggingface.co/datasets/ashritha0907/replay-gap-trajectories

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑