缺失的补全:面向编码智能体的状态条件最小充分证据
The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents
浏览论文内容
中文总结 AI 辅助
针对编码智能体,提出状态条件最小充分证据恢复方法MSS-Complement,将证据获取视为集合构建而非排名,在SERBench和AMA-Bench上显著提升决策支持完整性与准确率。
中文摘要 AI 辅助
一个在处理问题中途的编码智能体已经阅读了检索器排名最高的大部分内容。相关性是按段落评分的,但充分性属于集合:一个排名器可以用某个必需事实的变体填满其预算,却使决策缺乏支持。我们提出了状态条件最小充分证据恢复:给定捕获的智能体状态,恢复一个紧凑的证据组合,以提供其下一个决策仍缺乏的支持。SERBench在来自45个仓库的500个保留状态上对此进行衡量,记录智能体已看到的内容,并仅对覆盖当前决策注释要求的所有事实的集合给予信用。MSS-Complement将获取视为集合构建,而非排名。三次语义调用提出一个联合充分的集合,搜索其缺失的内容,并在6,144个令牌内返回4-8个完整的源代码单元。一个在校准数据上固定的配置,在五项时恢复了73.0%的这些状态的完整集合,在八项时恢复了80.6%,而Qwen3嵌入加重排分别为61.4%和72.4%。一个仅按相似性排序的匹配对照达到66.6%,将增益归因于集合级策略,而非计算。从冻结的仓库源代码且无金标准衍生池中,领先5.0个百分点。在AMA-Bench上,它从比该基准自身记忆智能体小76.2%的答案提示中回答,准确率高出2.08个百分点。从原本完整的集合中移除一个必需组,在两种执行器下分别使修复定位精度损失12.3和11.1个百分点。面向智能体的检索更应被表述为恢复决策所缺失的内容,而非对问题相似内容进行重排。
英文摘要
Retrieval assembles repository context by ranking passages for relevance to the current query. A coding agent halfway through an issue has already read much of what such a ranker returns. Relevance is scored per passage, but sufficiency belongs to the set: independently scored passages can fill the budget with support for one requirement while another goes unmet. We formulate state-conditioned minimal sufficient evidence recovery: given a captured agent state, recover a compact evidence combination supplying what its next decision still lacks. SERBench measures this on 500 held-out states from 45 repositories, recording what the agent has seen, crediting only sets that satisfy every annotated evidence requirement of the current decision, and separating set recovery from candidate discovery. MSS-Complement treats acquisition as set construction, not ranking. Three semantic calls propose a jointly sufficient set, search for what it lacks, and return 4-8 intact source units within 6,144 tokens. One configuration, fixed on calibration data, recovers a complete set for 73.0% of those states at five items and 80.6% at eight, against 61.4% and 72.4% for Qwen3 embedding with reranking. A matched control ranking by similarity alone recovers fewer complete sets, placing the margin over it in the set-level policy, not the computation. The lead persists from frozen repository source with no gold-derived pool. On AMA-Bench it answers from a 76.2% smaller answer prompt, with accuracy 2.08 points above that benchmark's own memory agent. Removing one required group from a complete set costs repair-localization precision under two executors. Retrieval for agents is better posed as recovering what a decision lacks than re-ranking what an issue resembles.
发表机构
- University of California San Diego(加州大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。