arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02267cs.AIcs.CLcs.CRcs.LG

快速模型,缓慢证据:LLM智能体框架中System-1决策模型的配对与自我审计评估

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

Jiawei Li

首次发表
浏览论文内容

中文总结 AI 辅助

本研究配对评估两个System-1决策模型在LLM智能体框架中的性能,发现托管模型更优,但审计揭示报告中的分析错误和混杂因素,强调证据验证的重要性。

中文摘要 AI 辅助

智能体框架在每个任务中需要做出许多小而类型化的决策:调用哪个模型、使用哪个工具、检索到的文本是否相关、输入是否携带注入。System-1决策模型通过单次前向传播和类别概率回答此类问题,相比LLM调用有望大幅节省成本和延迟。我们针对11个智能体决策点(构建自18个公共来源)对开源权重模型(Laya)和托管模型(Jev)进行了配对评估,包含7,283个基础案例及6,640个鲁棒性变体,采用字节相同输入、配对测试以及跨硬件和跨天可复现性检查。Jev在11个决策点中的9个上显著更准确(+10.8至+46.0个百分点)。两者在零样本模型路由上均未超过随机水平,在RAG相关性门控上持平。Laya在选项顺序反转时改变了30%的答案,并在候选众多或相似时性能急剧下降(50个最近邻工具时31%,而Jev在具有唯一正确工具的项目上为98%)。我们还审计了自身流程。三个分析错误和一个设计混杂因素扭曲了头条部署声明:遗漏的预筛选成本(报告节省23.9%,实际4.3%)、门控准确率被报告为端到端质量(58%对98%)、样本内阈值(目标5%,保留集缺失高达17%),以及一个“通道效应”在注入误报上消失于通道原生内容。另外两个疑似混杂因素未改变结论。所有案例、原始输出和分析代码可在该https URL获取。

英文摘要

Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp). Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating. Laya changes 30% of its answers when the option order is reversed and degrades sharply with many or similar candidates (31% at 50 nearest-neighbour tools, vs. 98% for Jev on items with a unique correct tool). We also audit our own pipeline. Three analysis errors and one design confound distorted headline deployment claims: an omitted pre-screen cost (reported 23.9% saving, actual 4.3%), gate accuracy reported as end-to-end quality (58% vs. 98%), in-sample thresholds (5% target, up to 17% held-out misses), and a "channel effect" on injection false positives that vanishes with channel-native content. Two other suspected confounds did not change the conclusions. All cases, raw outputs and analysis code are available at https://github.com/David-DL-Space/sys1-eval.

补充信息

↑