AI 中文总结
该研究扩展委托-代理框架至LLM委托场景,推导最优线性契约,用MATH、MMLUPro基准校准模型,证明线性契约可激励智能体的技术感知委托行为。
AI 中文摘要
我们将标准委托-代理框架扩展至代理人可从一组技术中进行选择的场景,每种技术具有不同的成本-能力特征。该框架在大语言模型(LLMs)时代愈发关键,此时代理人需同时选择模型及相关努力水平(如token预算)。我们将输出质量与努力的关系建模为凹型饱和函数,其取决于代理人隐藏的二维行动选择,该选择需平衡技术选择与努力分配。我们推导了委托人的最优线性契约,证明代理人的最优反应由触发技术切换的阈值报酬份额所表征。最后,我们使用MATH和MMLUPro基准上的开放权重LLM配对对模型进行校准,结果显示当委托人和代理人采用多臂老虎机算法在该环境中导航时,会收敛至与理论均衡高度一致的策略。这些结果表明,简单的线性契约可有效激励智能体工作流中复杂的、感知技术的委托行为。
英文摘要
We extend the standard Principal-Agent framework to scenarios where the Agent selects from a suite of technologies, each characterized by a distinct cost-capability profile. This framework is increasingly critical in the era of Large Language Models (LLMs), where Agents choose both a model and an associated effort level (e.g., token budget). We model the relationship between output quality and effort as a concave, saturating function, which depends on the Agent's hidden two-dimensional action choice balancing technology selection and effort allocation. We derive the optimal linear contract for the Principal, demonstrating that the Agent's best response is characterized by a threshold reward share that triggers technology switching. Finally, we calibrate our model using open-weight LLM pairings across the MATH and MMLUPro benchmarks. We show that both Principal and Agent, when employing bandit algorithms to navigate this environment, converge to strategies that closely align with our theoretical equilibrium. These results suggest that simple linear contracts can effectively incentivize complex, technology-aware delegation in agentic workflows.