发表机构
University of Florida; Aeronix(佛罗里达大学; Aeronix公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LatentSift利用策略隐藏状态进行无令牌无执行过滤,显著减少软件工程智能体验证令牌,同时保持或提升性能。
AI 中文摘要
测试时扩展通过生成多个候选轨迹并选择最佳轨迹来改进软件工程智能体。验证和选择这些长交互可能消耗与生成本身一样多的令牌。现有的混合工作流首先应用基于LLM的无执行(EF)验证器来过滤候选,然后再运行测试,这为每条轨迹增加了另一次模型传递。我们引入了LatentSift,一种无令牌且无执行的过滤器,用策略在生成候选时已经产生的隐藏状态取代了这一阶段。它通过推理、观察和函数调用状态来表示每个候选,将这些状态与策略训练期间从成功和不成功轨迹中收集的正负状态库进行比较,并将所得距离分数与学习的线性分数融合,以保留有希望的候选用于基于执行的阶段。在SWE-bench Verified上,跨三个智能体和两种策略规模,LatentSift在K=16时将EF验证器令牌削减了66.6–81.0%,并将包括测试生成在内的总验证令牌削减了49.1–62.1%,而混合Best@16匹配或改进了每个智能体的参考工作流,在DeepSWE-Preview上从59.26%提升到60.06%。
英文摘要
Test-time scaling improves software engineering agents by generating multiple candidate trajectories and selecting the best one. Verifying and selecting among these long interactions can consume as many tokens as generation itself. Existing hybrid workflows first apply an LLM-based execution-free (EF) verifier to filter candidates before running tests, which adds another model pass over every trajectory. We introduce LatentSift, a token-free and execution-free filter that replaces this first stage with hidden states the policy already produces while generating the candidates. It represents each candidate through its reasoning, observation, and function-call states, compares them with positive and negative banks of such states collected from successful and unsuccessful trajectories during policy training, and fuses the resulting distance scores with a learned linear score to retain promising candidates for the execution-based stages. On SWE-bench Verified, across three agents and two policy sizes, LatentSift cuts EF-verifier tokens by 66.6--81.0% and total verification tokens, which include test generation, by 49.1--62.1% at K=16, while hybrid Best@16 matches or improves on each agent's reference workflow, rising from 59.26% to 60.06% on DeepSWE-Preview.