AI 中文总结
研究广告商和用户委托代理的最优诚实描述问题,核心方法是基于代理的范围内遗憾,主要贡献是统一特定拍卖自动出价结果、确定语言模型诱导代理假设成立条件,揭示护栏三难困境及生产模型诚实报告存在未认领剩余等情况。
AI 中文摘要
广告商将出价委托给自动出价器;用户将任务委托给语言模型代理。人们向自动代理描述自己的需求,代理在机制中代表他们行事。这引出了经典理论未考虑的问题:何时向自动代理诚实地描述自己是最优的?我们表明答案取决于一个量,即代理的范围内遗憾。委托人通过误报所能获得的最大收益等于代理诚实报告行动相对于委托人本可引导其采取的行动的遗憾。当代理已经采取其能达到的最佳行动时,即忠诚时,诚实的自我描述是最优的(定理1)。该恒等式统一了特定拍卖的自动出价结果,并确定了语言模型诱导代理背后的忠实通信假设何时成立。该恒等式限制了对代理设置的护栏,从出价上限到模型的对齐层。没有一个护栏能同时具有约束力(它将真实行动从代理可达到的最佳结果中取代)、真实(诚实报告保持最优)和保持能力(该结果可通过某些报告达到);任意两者都会排除第三者(定理2)。一个改变模型行为同时使其最佳输出可达到的安全约束会使意图的诚实描述次优,因此更尖锐的报告可能会有收益。这就是提示工程和越狱背后的动机。由于精确计算范围内遗憾是#P难的,我们从样本中估计它,并在模型更新时维护它,成本由模型漂移的程度而非变化的频率决定。在对齐式上限下对五个提供商的生产语言模型运行该方法,我们发现诚实报告在每个模型上都留下了未被认领的剩余,通过夸大报告可以恢复。
英文摘要
Advertisers delegate bidding to autobidders; users delegate tasks to language-model agents. A person describes what they want to an automated proxy that acts in a mechanism on their behalf. This is the revelation principle in production, and it forces a question classical theory assumes away: when is it optimal to describe yourself honestly to your own proxy? We show the answer turns on one quantity, the proxy's within-range regret. The most a principal can gain by misreporting equals the regret of the proxy's honest-report action against those the principal could have steered it to take. Honest self-description is optimal exactly when the proxy already plays the best action it can reach, that is, when it is loyal (Theorem 1). The identity unifies auction-specific autobidding results and pins down when the faithful-communication assumption behind language-model elicitation proxies (Huang et al.) holds. The identity constrains guardrails placed on proxies, from bid caps to a model's alignment layer. No guardrail can be at once binding (it displaces the truthful action from the proxy's best reachable outcome), truthful (honest reporting stays optimal), and capability-preserving (that outcome stays reachable through some report); any two preclude the third (Theorem 2). A safety constraint that alters what a model does while leaving its best output reachable makes honest description of intent suboptimal, so a sharper report can gain. This is the incentive behind prompt-engineering and jailbreaking. Because within-range regret is #P-hard to compute exactly, we estimate it from samples and maintain it as a model is updated, at a cost set by how far the model drifts, not how often it changes. Running it on production language models from five providers under an alignment-style cap, we find honest reporting leaves surplus unclaimed on every model, recovered by inflating the report.
Comments31 pages, 7 figures, 4 tables