发表机构
University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控实验揭示,在AI生产中,工作流与信息结构比资源规模更关键,任务知情规划在高资源下优势显著,直接执行则更高效。
AI 中文摘要
推理的经济价值取决于能力和任务信息在AI生产各阶段之间的分布方式。我们通过在外部验证的软件工程任务上进行的受控工作流实验来研究这些组织边际。在两个匹配的资源面板中,直接执行在逻辑令牌上限为12,000和24,000时记录了相同的59.6%成功率,而信息受限规划下的成功率从36.2%上升至51.2%。规划的劣势缩小了15.0个百分点(95%任务簇自助法区间:4.2至25.8)。一项严格的只读规划活动改变了规划者是否看到任务问题。在12,000令牌时,问题访问使成功率比隐藏问题的规划提高了约16个百分点。与直接执行相比,任务知情规划在12,000令牌时低约10个百分点;在24,000令牌时,它显示出29.6个百分点的优势。在资源面板中,直接执行使用的资源远低于任一上限,而规划工作流的绑定率从46.2%降至0.8%,下游执行占总使用量增加的89.9%。规模决定了系统可用的容量;工作流和信息结构塑造了生产价值。
英文摘要
The economic value of inference depends on how capacity and task information are distributed across stages of AI production. We study these organizational margins using controlled workflow experiments on externally verified software-engineering tasks. In two matched resource panels, direct execution records the same success rate of 59.6 percent at logical-token ceilings of 12,000 and 24,000, while success under information-constrained planning rises from 36.2 to 51.2 percent. The planning disadvantage narrows by 15.0 percentage points (95 percent task-cluster bootstrap interval: 4.2 to 25.8). A strict read-only planning campaign varies whether the planner sees the task issue. At 12,000 tokens, issue access raises success by about 16 percentage points over issue-hidden planning. Compared with direct execution, task-informed planning is about 10 points lower at 12,000 tokens; at 24,000 tokens, it shows a 29.6-point advantage. In the resource panels, direct execution uses substantially less than either ceiling, while the planning workflow's binding rate falls from 46.2 to 0.8 percent and downstream execution accounts for 89.9 percent of the increase in total use. Scale determines the capacity available to a system; workflow and information structure shape the productive value