arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17614cs.MAcs.SYeess.SY

基于核化多臂老虎机的动态委托代理问题中的自适应激励设计

Adaptive Incentive Design in Dynamic Principal-Agent Problem via Kernelized Bandits

Arghya Mallick, Anuj S. Vora, Sergio Grammatico, Peyman Mohajerin Esfahani

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对动态委托代理问题的现有局限,提出引入随机效用模型的Heteroscedastic GP-UCB算法,建立了累积遗憾界,并在V2G激励设计中验证了其更优的经济性能。

中文摘要 AI 辅助

我们考虑信息不对称下的动态委托代理问题,其中委托方需依次设计合约,以激励具有未知偏好和隐藏行动的代理方。现有文献的一个基本瓶颈是假设代理方效用具有确定性,这会导致委托方的期望效用不连续,并迫使对合约空间进行计算上难以处理的离散化。在本文中,我们通过在代理方效用模型中引入随机对应项来解决这一局限,该对应项可捕捉现实子系统中固有的物理和行为变化。我们正式证明,这种随机形式恢复了委托方期望效用的连续性。利用这种连续几何结构,我们将交互过程建模为受异方差噪声约束的结构化多臂老虎机问题。我们提出了一种\texttt{Heteroscedastic GP-UCB}算法,该算法利用神经网络(Arcsin)核,此核被选用于捕捉效用景观的非平稳、S型几何特征。对于$m$维紧凑合约空间,我们建立了高概率下的累积遗憾界为$O\left(\sqrt{T}(\log T)^{m+1}\right)$。最后,我们通过构建车到网(Vehicle-to-Grid, V2G)激励设计问题,证明其与动态委托代理问题等价,并展示其对电网聚合商而言具有更优的经济性能,以此验证我们理论框架的实际效用。

英文摘要

We consider the dynamic principal-agent problem under asymmetric information, wherein a principal sequentially designs contracts to incentivize an agent with unknown preferences and hidden actions. A fundamental bottleneck in the existing literature is the assumption of deterministic agent utility, which renders the principal's expected utility discontinuous and forces computationally intractable discretizations of the contract space. In this paper, we address this limitation by introducing a stochastic counterpart into the agent's utility model, capturing the inherent physical and behavioral variations in realistic subsystems. We formally prove that this stochastic formulation restores the continuity of the principal's expected utility. Leveraging this continuous geometric structure, we formulate the interaction as a structured multi-armed bandit problem subject to heteroscedastic noise. We propose a \texttt{Heteroscedastic GP-UCB} algorithm that utilizes a Neural Network (Arcsin) kernel, chosen to capture the non-stationary, sigmoidal geometry of the utility landscape. For an $m$-dimensional compact contract space, we establish a high-probability cumulative regret bound of $O\left(\sqrt{T}(\log T)^{m+1}\right)$. Finally, we demonstrate the practical efficacy of our theoretical framework by formulating the Vehicle-to-Grid (V2G) incentive design problem, proving its equivalence to a dynamic principal-agent problem, and showing superior economic performance for grid aggregators.

↑