具有无限制有界合约的极小极大最优在线合约设计
Minimax-Optimal Online Contract Design with Unrestricted Bounded Contracts
浏览论文内容
中文总结 AI 辅助
研究重复合约设计中委托人仅观察结果时的在线学习问题,提出基于有效维数约简与Lipschitz参数化的策略,实现极小极大遗憾阶T^{m/(m+1)},并证明每个额外可合约结果都会精确增加学习成本。
中文摘要 AI 辅助
我们研究了当委托人观察结果但无法观察产生这些结果的行动时的重复合约设计问题。委托人可以使用任何有界的结果依赖支付向量,而代理人的最优反应可能使预期利润在这些支付中不连续。对于每个固定的结果数量 $m\ge2$,在 $T$ 轮上的极小极大遗憾的阶为 $T^{m/(m+1)}$(忽略对数因子)。上界允许任意的行动空间和代理异质性,且无需平滑性或单调剩余假设。其关键在于一个有效维数约简:即使固定平局打破规则不是平移不变的,基准也可以被归一化,之后显示性偏好产生一个在支付差坐标下的单调反应映射。基于该映射的Lipschitz参数化构建的学习策略仅使用观察到的结果类别即可达到该速率。下界构造考虑了激励损失如何在结果维度上累积。它表明每个额外的可合约化结果都会导致学习的最坏情况成本产生精确且不可避免的增加。
英文摘要
We study repeated contract design when a principal observes outcomes but not the actions that generate them. The principal may use any bounded outcome-contingent payment vector, and the agent's best response can make expected profit discontinuous in those payments. For every fixed number $m\ge2$ of outcomes, the minimax regret over $T$ rounds is of order $T^{m/(m+1)}$, up to logarithmic factors. The upper bound allows arbitrary action spaces and agent heterogeneity, without smoothness or monotone-surplus assumptions. Its key is an effective-dimension reduction that the benchmark can be normalized even when fixed tie-breaking is not shift invariant, after which revealed preference yields a monotone response map in payment-difference coordinates. A learning policy built on a Lipschitz parametrization of this map attains the rate using only observed outcome categories. The lower-bound construction accounts for how incentive losses accumulate across outcome dimensions. It shows that each additional contractible outcome creates a precise and unavoidable increase in the worst-case cost of learning.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。