发表机构
Pheo Inc(Pheo 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体重复任务不一致和浪费的问题,提出技能习惯形成机制,通过确定性技能竞争与四门槛验证,在文本到SQL任务中实现完全复现并减少令牌使用,提升准确性与可审计性。
AI 中文摘要
在重复工作中,智能体的表现不一致。我们将42项任务各运行三次后发现,根据模型不同,38%至74%的返回答案不一致。一致性是买家、审计员或监管机构所要求的,而智能体并不具备。它们也很浪费:智能体生成的内容中有95.3%至97.2%用于重新推导系统已知的计划。我们提出技能习惯形成机制。智能体从自身的执行历史中挖掘候选技能,这些确定性变体与现有技能竞争而非取代它。候选技能声明其声称的输入空间区域,因此常见情况按脚本运行,其余情况则回退到推理。四个成本递增的门槛接纳候选技能;核心门槛将候选技能的执行轨迹与保留的参考轨迹进行对比,容差基于该参考轨迹自身的运行间变异性来衡量。在文本到SQL任务中,四个推理分支中有三个在42个重复问题中的11至13个上复现了自身输出,第四个在26个上复现,而一个习惯形成的变体在我们重复的所有456次调度中全部复现,并且不劣于它所取代的每个分支(p<0.0001)。它还使用了少14%至56%的令牌,在7至53次重用后转为净收益。我们衡量了这在准确性上的代价。守卫在2.6%的自然改写和26%的边界附近输入上接纳了本应推迟的工作,且13个此类失败中有11个在任何阈值下对轨迹一致性门不可见。确定性错误会精确重复:坏习惯和好习惯一样可靠,这正是使系统可审计的属性所付出的代价。将路由与参数提取分离,在成本为43%的情况下,端到端准确性从0.888提升到0.952。
英文摘要
On repeated work, agents are inconsistent. We ran 42 tasks three times each and found that, depending on the model, 38% to 74% returned answers that did not agree. Consistency is what a buyer, an auditor, or a regulator requires, and agents do not have it. They are wasteful too: 95.3% to 97.2% of what an agent generates goes to re-deriving a plan the system already knows. We propose skill habit formation. An agent mines its own execution history for candidate skills, deterministic variants that compete against the incumbent rather than replacing it. A candidate declares the region of input space it claims, so the common case runs as a script and the rest falls through to reasoning. Four gates of ascending cost admit candidates; the central one tests a candidate's execution trace against a retained reference, within a tolerance measured from that reference's own run-to-run variability. On text-to-SQL, three of four reasoning arms reproduced their own output on 11 to 13 of 42 repeated questions and the fourth on 26 of 42, while a habit-formed variant reproduced on all 456 dispatches we repeated and was non-inferior to every arm it replaced (p<0.0001). It also used 14% to 56% fewer tokens, turning net positive after 7 to 53 reuses. We measured what this costs in accuracy. The guard admitted work it should have deferred on 2.6% of natural paraphrases and 26% of inputs near its boundary, and 11 of 13 such failures were invisible to the trace-conformance gate at any threshold. Deterministic errors repeat exactly: a bad habit is as reliable as a good one, and that is the price of the property that makes the system auditable. Separating routing from parameter extraction raised end-to-end accuracy from 0.888 to 0.952 at 43% of the cost.
Comments12 pages, 2 figures