AI 中文总结
该研究针对多调用LLM工作流的模型替换问题,通过谓词-动作分解分离决策,经多任务实验揭示模型价值依赖角色与部署,提出大规模工作流路由的实用序列。
AI 中文摘要
大型多调用大语言模型(LLM)系统存在一个查询级路由无法捕捉的科学问题:模型的价值取决于其在依赖计算中的入口位置,以及该调用所处的部署环境。现有路由算法通常仅决定“在哪里使用更强的模型”,却将替换本身的收益视为已知。我们通过谓词-动作分解将这两个决策分离,并在受控的求解-合并-验证工作流中进行评估,该工作流涵盖8至64次求解调用,以及四个三级模型阶梯。所得证据揭示了看似冲突结果背后的一致原则:在数值频率计数任务中,混合通义千问(Qwen)/GPT阶梯的全强模型配置将均方根误差(RMSE)从4.818降至1.538,而仅使用通义千问的平均排序则出现反转;输入匹配干预进一步表明,相同的中等到强的动作在不同角色和规模下具有截然不同的价值;语义任务与契约转变再次反转了混合排序,而分配消融实验则区分出有用的稀疏放置与覆盖不足及无差别升级的差异。这些结果共同确立了模型替换是一种受部署条件约束的动作,而非由层级标签隐含的属性,并为大规模工作流路由提供了实用序列:校准动作、解析其角色依赖的效果,随后优化其放置。
英文摘要
Large multi-call LLM systems pose a scientific problem that query-level routing does not capture: the value of a model depends on where it enters a dependent computation and on the deployment that surrounds that call. Existing routers typically decide \emph{where} to spend a stronger model while treating the benefit of the substitution itself as known. We separate these two decisions through a predicate-action factorization and evaluate it in controlled solve-merge-verify workflows spanning 8-64 solve calls and four three-tier model ladders. The resulting evidence reveals a consistent principle beneath apparently conflicting outcomes. On numeric frequency counting, all-strong reduces RMSE from 4.818 to 1.538 in the Mixed Qwen/GPT ladder, whereas the average Qwen-only ordering reverses. Input-matched interventions further show that the same medium-to-strong action has sharply different value across roles and scales. A semantic task-and-contract shift reverses the Mixed ordering again, while allocation ablations distinguish useful sparse placement from under-coverage and indiscriminate escalation. Together, these results establish model substitution as a deployment-conditioned action rather than a property implied by a tier label, and they provide a practical sequence for large-scale workflow routing: calibrate the action, resolve its role-conditioned effect, and then optimize its placement.