你能检查吗?本地LLM网络自动化的可检查性边界
Can You Check That? The Checkability Boundary for Local LLM Network Automation
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出可检查性标准,通过Touchstone流水线利用本地小模型和内在检查,在冲突检测和意图翻译任务上达到高精度并减少升级,验证了本地推理的适用边界。
AI中文摘要:
将每个网络自动化输入发送给第三方前沿LLM会导出敏感工件,如生产配置、拓扑和日志。在本地查询小型语言模型(SLMs)可避免这种外泄,但SLM的输出可能容易出错,不适合直接使用。本工作引入可检查性作为确定哪些任务适合本地推理的标准。当任务暴露一个廉价的确定性测试——内在检查——拒绝违反必要正确性条件的输出时,该任务是可检查的。我们在Touchstone中实例化这一想法,这是一个本地优先的流水线,使用七个现成的SLM(1-8B参数)生成候选,使用任务特定的内在检查拒绝响应,并将未解决的输入升级到前沿LLM。在冲突检测和意图翻译任务上,Touchstone分别达到98.6%和93.8%的端到端准确性,同时仅升级16%和17%的输入。在TeleQnA上,一个仅知识控制且没有任务特定内在检查的任务,Touchstone无法匹配前沿基线的准确性。我们的结果支持一个简单的部署规则:当任务语义支持精确、低成本的检查时,保持推理本地化;其余升级。
英文摘要:
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating a necessary correctness condition. We instantiate this idea in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a frontier LLM. On conflict detection and intent translation tasks, Touchstone reaches 98.6% and 93.8% end-to-end accuracy while escalating only 16% and 17% of inputs, respectively. On TeleQnA, a knowledge-only control that has no task-specific intrinsic checks, Touchstone is unable to match the accuracy of the frontier baseline. Our results support a simple deployment rule: keep inference local when task semantics support precise, low-cost checks; escalate the rest.