发表机构
PayPal AI(PayPal AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建微任务基准,发现现成小语言模型在智能体框架中普遍未达资格阈值,量化无改善,建议将其置于合格基线之后使用。
AI 中文摘要
智能体框架越来越多地希望在前沿大语言模型(LLM)规划器周围的微任务上运行小语言模型(SLM):自动批准shell命令、编写记忆、选择工具、对过去的轮次进行排序。我们询问现成的SLM是否满足从业者定义的阈值,以及当它们失败时,原因是什么,以及量化是否会改变答案。我们构建了一个包含4个此类微任务的基准,使用固定提示和自动指标,每个任务都有一个预先指定的阈值τ,该阈值锚定于一个廉价的非LLM基线,并采用CI感知的资格规则(仅当配置的置信区间下界超过τ时,该配置才通过)。在Qwen3 0.6/1.7/4/8B的最佳设置(FP16、贪婪解码、一个冻结提示、无调优)下进行扫描,我们发现了一个资格差距:16个配置(4个任务×4个模型)中0个通过(通过检查原始输出和解析器行为验证)。一个对数概率决策阈值诊断(T1/T3/T4;T2通过上下文长度/级联探针)将失败区分为能力缺陷和可以通过改变解码阈值来解决的失败(4种机制)。量化为4位(RTN/GPTQ/AWQ)会造成取决于模型大小的损害,并且不会使任何配置进入合格状态(在可重建的硬标签任务T1/T3上认证,在T2/T4上进行诊断/窗口稳健性),因此差距更多地取决于模型大小而非精度;它在Llama-3.x上复现(12/12不合格),并且对锚点选择(τ扫描)和提示措辞(原始加上每个单元格3个中性改写,共112个配置中0个合格)具有稳健性。实际意义:将SLM置于满足CI支持阈值的基线之后,并且仅在基线未能达到阈值时使用SLM;例如,在BM25候选列表上的4B重排序器优于BM25(+0.047 [0.020, 0.073],而本身并不证明合格性)。
英文摘要
Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell commands, writing memory, selecting tools, ranking past turns. We ask whether off-the-shelf SLMs meet practitioner-defined thresholds and, when they fail, why, and whether quantization changes the answer. We build a benchmark of 4 such microtasks with fixed prompts and automatic metrics, each with a pre-specified threshold $τ$ anchored to a cheap non-LLM baseline and a CI-aware eligibility rule (a configuration passes only if its confidence bound clears $τ$). Sweeping Qwen3 0.6/1.7/4/8B at their best (FP16, greedy, one frozen prompt, no tuning), we find an eligibility gap: 0 of 16 (4 tasks $\times$ 4 models) configurations pass (verified by checking the raw outputs and parser behavior). A logprob decision-threshold diagnostic (T1/T3/T4; T2 via a context-length/cascade probe) separates the failures into capability deficits and failures that can be addressed by changing the decoding threshold (4 regimes). Quantization to 4-bit (RTN/GPTQ/AWQ) does damage that depends on model size and moves no configuration into eligibility (certified on the reconstructable hard-label tasks T1/T3, diagnostic/windowed robustness on T2/T4), so the gap tracks model size more than precision; it replicates on Llama-3.x (12/12 ineligible) and is robust to the anchor choice (a $τ$-sweep) and to prompt wording (0/112 eligible across the original plus 3 neutral paraphrases per cell). The practical implication: place SLMs behind a baseline that meets the CI-backed threshold, and use the SLM only where the baseline fails to meet the threshold; e.g. a 4B re-ranker over a BM25 shortlist beats BM25 ($+0.047$ [0.020, 0.073], without itself certifying eligibility).
CommentsPreprint. under review at a NeurIPS 2026 workshop. 15 pages, 8 figures, 14 tables