arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.05813cs.GTcs.LG

带强盗学习的上下文采购拍卖

Contextual Procurement Auctions with Bandit Learning

Yiling Chen, Shi Feng, Sadie Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

研究重复上下文采购拍卖中平台从强盗反馈学习依赖上下文产品价值的问题,给出有\(\widetilde O((ng)^{1/3}T^{2/3})\)遗憾值的先探索后提交机制及固定支付UCB机制,证明了固定支付遗憾值与激励权衡的匹配下界。

中文摘要 AI 辅助

我们研究了重复的上下文采购拍卖,其中平台必须从强盗反馈中学习依赖上下文的产品价值。我们给出了一种具有\(\widetilde O((ng)^{1/3}T^{2/3})\)遗憾值的完全真实的先探索后提交机制。我们还给出了一种具有遗憾值与激励权衡的固定支付UCB机制:接近UCB的调整实现了\(\widetilde O(\sqrt{ngT})\)的福利遗憾值,而对于固定的\(n,g\),其总激励误差为\(\widetilde O(T^{3/4})\);平衡调整在两个尺度上都给出了\(\widetilde O(T^{2/3})\)。遗憾值是相对于全信息有效分配的福利损失来衡量的。我们证明了固定支付遗憾值与激励权衡的匹配下界。

英文摘要

We study repeated procurement auctions in which producers have private costs and the platform must learn the context-dependent value of selecting each producer. We evaluate performance by welfare regret: the cumulative loss in total surplus relative to the full-information efficient rule that knows the context-dependent values and true producer costs. The natural UCB allocation rule achieves $\widetilde O(\sqrt{ngT})$ welfare regret under truthful bids, but its adaptive, bid-dependent learning path does not by itself ensure truthfulness. To obtain exact incentives, we first design a bid-independent explore-then-commit mechanism with empirical threshold payments; it is dominant-strategy truthful and has $\widetilde O((ng)^{1/3}T^{2/3})$ regret. We then introduce frozen-payment UCB, which estimates payments from initial bid-independent exploration but continues allocation learning by UCB. Under a truthful-path margin condition, the frozen-payment UCB is approximately truthful with an average per-round deviation gain $\widetilde O(T^{-1/4})$ for fixed $n$, $g$. Under truthful bidding, it achieves $\widetilde O(\sqrt{ngT})$ welfare regret, matching the UCB rate. A lower bound shows that this regret-incentive tradeoff is tight within the frozen critical-payment class.

补充信息

↑