什么阻碍了小语言模型驱动数据库智能体
What Stops a Small Language Model From Driving a Database Agent
浏览论文内容
中文总结 AI 辅助
本研究在生产系统中测试39个小型开放权重语言模型驱动SQL智能体的能力,发现75.7%的失败源于已调用工具但未获结果的运输问题,并揭示服务器缺陷及上下文窗口混杂因素。
中文摘要 AI 辅助
小型开放权重语言模型被认为在智能体数据库工作中表现不佳,因为它们缺乏相应的推理能力。我们针对一个生产系统对此进行了测试。在十一天内,我们使用39个本地服务的开放权重模型和一个托管控制模型,驱动了一个开源SQL客户端的智能体模式,覆盖六个任务面:共进行8,199次运行,产生110,711条账本事件,以及14,008次被拒绝的工具调用。在2,100次归因于模型的智能体模式损失中,有1,590次(占75.7%)来自至少调用过一次工具的运行。这一多数情况在重采样模型而非运行后依然成立:在99.7%的聚类重采样中以及22个至少有20次损失的模型中的15个模型中均保持。其中,运输失败(即使用了工具但从未获得可交付结果的运行)是最大类别,占36.2%,而能力失败(即未调用任何工具的运行)是最小类别,占17.3%;我们将这一排序报告为该语料库的属性而非普遍发现,因为按模型聚类后,该排序仅在74.5%的重采样中成立。运输失败可分解为几种机械性的参数形态。生产账本记录拒绝代码,但从不记录模型的参数,因此这些在十天内不可见;捕获它们暴露了五个服务器缺陷,其中一个缺陷要求在一个工具上设置某个字段,禁止在其组成工具的兄弟工具上设置该字段,然后因该字段缺失而判定运行失败。五次服务器更改(不涉及模型、提示词或采样设置)使六个模型在30个单元格中移动了6至21个单元格。我们还报告了一个我们认为影响已发布本地模型基准(包括我们自己的)的混杂因素:在没有上下文上限的情况下,一个7.1 GB的模型以其完整的262,144令牌窗口被接纳,并在64 GB机器上占用51 GB内存,产生在普通日志中与模型超时无法区分的运行。语料库、评分器以及一个可重新生成每个图形的验证器均已发布。
英文摘要
Small open-weight language models are assumed to fail at agentic database work because they lack the reasoning capacity for it. We test that against a production system. Over eleven days we drove the agent mode of an open-source SQL client with 39 open-weight models served locally and one hosted control, across six task surfaces: 8,199 runs, 110,711 ledger events, 14,008 refused tool calls. Of the 2,100 model-attributed agent-mode losses, 1,590, or 75.7%, came from runs that had invoked at least one tool. That majority is what survives resampling models rather than runs: it holds in 99.7% of clustered resamples and in 15 of the 22 models with at least twenty losses. Within it, transport, a run that used the tools and never got a deliverable through, is the largest class at 36.2% and capability, a run that invoked no tool at all, the smallest at 17.3%; we report that ordering as a property of this corpus rather than a general finding, since clustered by model it holds in only 74.5% of resamples. Transport failures decompose into a few mechanical argument shapes. Production ledgers record refusal codes and never the model's arguments, so these were invisible for ten days; capturing them exposed five server defects, one of which demanded a field on one tool, forbade it on the sibling that composed it, then failed the run for its absence. Five server changes, touching no model, prompt or sampling setting, moved six models by 6 to 21 cells out of 30. We also report a confound we believe affects published local-model benchmarks, ours included: with no context cap, one 7.1 GB model was admitted at its full 262,144-token window and held 51 GB on a 64 GB machine, producing runs indistinguishable in any ordinary log from a model timing out. The corpus, the scorer and a verifier that regenerates every figure are released.