arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UnitBoost:用合并算子而非模型管理复合LLM系统

UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

Xing Zhang, Guanghui Wang, Yanwei Cui, Mengdie Flora Wang, Peiyang He

arXiv 2609.09815首次发表:更新:

发表机构

AWS Generative AI Innovation Center(AWS生成式AI创新中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UnitBoost用元级合并算子取代高层LLM管理复合系统,实现顺序无关、带单元来源和可测试失败条件,在多个基准上显著提升性能。

AI 中文摘要

复合LLM系统通常通过添加一个更高级别的LLM来解决协调问题。由此产生的元智能体读取各工作单元的输出,撰写最终答案,分配后续调用,并决定何时停止。这种方法具有表达力,但也将三个控制决策集中在一次不透明、对顺序敏感的模型调用中。我们质疑管理者是否必须具有生成性。UnitBoost用一个定义明确的元级算子取代了该模型:一个任务给定的单元映射将工作单元输出转化为槽值提议,一个受约束的argmax组装输出,而未被填充或未被支持的槽位则成为下一轮的显式残差。该算子与顺序无关,记录单元来源,并给出一个简单保证:在没有耦合约束的情况下,在相同准入分数下进行单元级最大化支配对任何完整候选的选择。在三个保留基准上,它比使用黄金标签选择的最佳单一候选高出0.060-0.195个绝对任务分数点,比输入匹配的生成式管理者高出0.048-0.076。仅替换管理步骤即可改善六个复合系统配置,提升幅度为0.013-0.182。残差导向轮次将FanOutQA单元格F1从0.4778提升至0.5524;匹配对照表明,真实残差优于随机目标和普通重读,而无标签供给信号在一轮无产出后标记耗尽。同样的分析衡量了三种无法获得此类增益的条件(一个不可分割单元、单元身份不可用以及每个发出单元都收费的端点),并将跨单元耦合量化为修复成本。管理者放弃语义自由度,获得顺序不变性、单元来源和可测试的失败条件。

英文摘要

Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting meta-agent reads workers' outputs, writes the final answer, allocates later calls, and decides when to stop. It is expressive, but it also concentrates three control decisions in an opaque, order-sensitive model call. We ask whether the manager needs to be generative at all. UnitBoost replaces that model with a defined meta-level operator: a task-given unit map turns worker outputs into slot-value proposals, a constrained argmax assembles the output, and the slots left unfilled or unsupported become an explicit residual for the next round. The operator is order-free, records unit provenance, and gives a simple guarantee: without coupling constraints, unit-wise maximization under the same admission score dominates selection of any complete candidate. On three held-out benchmarks, it exceeds the best single candidate chosen with gold labels by 0.060-0.195 absolute task-score points and input-matched generative managers by 0.048-0.076. Replacing only the management step improves six compound-system configurations by 0.013-0.182. Residual-directed rounds raise FanOutQA cell F1 from 0.4778 to 0.5524; matched controls show that the true residual outperforms random targets and ordinary rereading, while a label-free supply signal flags exhaustion after one unproductive round. The same analysis measures three conditions in which no such gain is available (one indivisible unit, unavailable unit identity, and an endpoint that charges for every emitted unit) and quantifies cross-unit coupling as a repair cost. The manager gives up semantic freedom and gains order invariance, unit provenance, and testable failure conditions.

CommentsAccepted at the NeurIPS 2026 Workshop on Managing Agents that Manage Agents

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑