Mingbird:一种本地优先的智能体框架,使小型开放模型能够完成真实任务
Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks
浏览论文内容
中文总结 AI 辅助
针对小型开放模型在云端框架下难以完成真实任务的问题,提出本地优先的Mingbird框架,通过十种机制逐点补偿模型缺陷,在LRAB和τ²-bench基准上显著提升任务完成得分。
中文摘要 AI 辅助
小型开放权重模型(2-9B)可在普通笔记本电脑上运行,但在云端规模的智能体框架下,它们很少能完成真实任务:工具预填充溢出上下文,自我纠正发生偏离,工具演示陷入循环,任务被静默放弃。我们通过受控的单机比较和一个第三方基准提供证据,表明这些失败中有相当一部分归因于框架而非模型。我们介绍了Mingbird,一个面向Windows和Ollama的本地优先智能体框架,其十种机制逐点补偿小型模型的失败形式,其中三种具有代表性:字节级净零预填充预算、在接受完成前重新读取任务的完成门控,以及签名级循环检测。在LRAB上,一个保持机器、模型、预算和评分固定的受控比较(4个框架 × 4个开放模型(2B-35B) × 18个真实任务,确定性工件评分),Mingbird达到0.886的总体得分,而goose为0.631,opencode为0.479,agent-mini为0.405,所有288个单元均已发布;在τ²-bench(278个任务,三个臂,一个协议)上,其总得分为0.856,对比0.791和0.737;对相同18个任务的前沿模型探针在框架间跨度从0.997到0.478,而结构良好的脚手架彼此之间保持在0.072以内。留一机制消融仅作为方向性报告:同一臂的当晚重复实验使其均值移动高达0.069,相当于每个名义单次试验的增量大小,而一次批次匹配的比较(完整机制栈对比仅文本重读)在三次重复中为可执行完成门控提供了配对+0.10的增益。证据带有明确的局限性:自建基准、单台机器和单次试验评分。
英文摘要
Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We introduce Mingbird, a local-first agent harness for Windows and Ollama whose ten mechanisms compensate point-by-point for small-model failure forms, three of them representative: a byte-level net-zero prefill budget, a finish gate that re-reads the task before accepting completion, and signature-level loop detection. On LRAB, a controlled comparison holding machine, models, budgets, and scoring fixed (4 harnesses $\times$ 4 open models (2B-35B) $\times$ 18 real tasks, deterministic artifact scoring), Mingbird reaches 0.886 overall against 0.631 (goose), 0.479 (opencode), and 0.405 (agent-mini), with all 288 cells published; on $τ^2$-bench (278 tasks, three arms, one protocol) it totals 0.856 against 0.791 and 0.737; and a frontier-model probe on the same 18 tasks spans 0.997 to 0.478 across harnesses, with well-formed scaffolds staying within 0.072 of each other. A leave-one-mechanism-out ablation is reported as directional only: same-night replications of the same arm move its mean by up to 0.069, the size of every nominal single-trial delta, and the one batch-matched comparison (full mechanism stack versus text re-read alone) gives the executable completion guards a paired +0.10 across three replications. The evidence carries stated limits: a self-built benchmark, a single machine, and single-trial scoring.
发表机构
- University of Science and Technology Beijing(北京科技大学)
- Honor Device Co., Ltd.(荣耀终端有限公司)
机构由 AI 辅助整理,请以论文原文为准。