arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

完成的配对掩盖了封顶失败:选择性上下文投影的 ReVerPi 案例研究

Evaluating Budgeted Context Projection with Unexecuted Companion Runs

Guangzhe Zhang

arXiv 2609.31381首次发表:更新:

发表机构

Independent AI Researcher(独立人工智能研究员)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过 ReVerPi 案例,分析上下文投影在减少重复输入与增加检索轮次间的权衡,发现完成配对掩盖边界失败,并建议评估保留干预边界、独立执行双臂并报告资源消耗。

AI 中文摘要

上下文投影用紧凑、可寻址的摘录替换较旧的工具观测结果,减少重复输入,同时可能增加证据检索轮次。我们在 ReVerPi(一个具有归档观测结果和匹配的完整/投影连续性的 Pi 扩展)中研究这一权衡。在一项包含 641 次模型请求的 86 次运行源码阅读活动中,15 个完成的配对显示出相同的成功率:每臂 12/15。另有 12 次边界运行停止,当第一臂未能完成时,运行器抑制了伴随臂。恢复所有 27 次边界运行将投影减去完整的成功率限制在 -9 到 +1 个任务之间。一次被省略的、由选择器选择的投影连续性成功检索到归档文本,但耗尽了十二次请求;其完整对应物在三次内回答。十一个共同正确的配对在此记录框架内形成一个完全观测的成功层:投影将聚合逻辑令牌减少 25%,同时将配对的中位令牌增加 29%,并将总后缀请求从 35 增加到 55。将拟合与评估分开改变了选择器的表面平局:在其四个拟合配对之外,它在十三次可比运行中多 incur 一次失败和 8.6% 的逻辑令牌。这项方法论案例研究将停止规则、已知的有界失败、未执行的伴随臂和资源聚合联系起来。其发现涉及记录的活动,而非总体非劣性或优于无限制 Pi。评估应保留每个干预边界,独立于第一臂的完成执行两个分配的臂,并报告完成情况以及交互和令牌消耗。

英文摘要

Context projection can shorten individual requests while changing whether an agent finishes within its budget. We examine how sequential evaluation obscures this trade-off when a capped first continuation prevents its companion from running. In a recorded ReVerPi source-reading campaign, 15 pairs with two final answers yield 12 historically scored successes per arm. Retaining all 27 intervention boundaries distinguishes observed failures from ten unexecuted companions and bounds projected-minus-full success between $-9$ and $+1$ tasks. Under the archived scoring contract, a frozen projection selector has a success difference from full context of $[-3,0]$; outside four fitting tasks, it is $[-4,-1]$ across 23 boundaries. Excluding one task whose platform premise is not established by retained actor-input evidence changes the nonfitting range to $[-3,0]$ across 22 boundaries, so strict inferiority is not robust to that exclusion. Among eleven historically joint-success pairs, projection uses 25% fewer aggregate logical tokens but more tokens for the median pair and 55 rather than 35 suffix requests. These retrospective results concern one adaptively assembled campaign, not population performance. The case motivates accounting that retains every boundary, preserves unknown outcomes and policy dependencies, checks task premises, and separates bounded completion from success-conditioned resource use.

Comments23 pages. Revised analysis and presentation; expanded sensitivity analyses and offline reproducibility materials. Project code: https://github.com/timwhitez/ReVer_Pi. Source archive includes ancillary data and offline reanalysis scripts

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑