arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIRT:用于具有全信息流排列外部性的真实生成式拍卖的Transformer

MIRT: Transformers for Truthful Generative Auctions with Whole-feed Permutation Externalities

Ali Elahi, Ermis Soumalias, Jason Cheuk Nam Liang, Daniel Yao, Michael J. Curry

arXiv 2610.06559首次发表:更新:

发表机构

University of Illinois Chicago; Meta(伊利诺伊大学芝加哥分校; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MIRT机制,用Transformer生成广告与自然内容联合排序的候选信息流并选福利最大者,通过强化学习实现出价无关,兼顾策略proofness与外部性优化。

AI 中文摘要

现代在线平台通常先分别对广告和自然内容进行排序,再将其混合成展示给用户的信息流,这忽略了外部性:一个条目的点击率取决于其周围内容,而不仅仅取决于其自身位置。最近的基于学习的生成式信息流机制对部分此类跨类型交互进行建模,以全局优化整个信息流的福利。然而,这些方法要么固定自然内容的顺序,要么缺乏对竞拍者的精确策略proofness保证。为克服这些不足,我们引入了最大范围内Transformer(MIRT)机制类,它使用Transformer生成一个候选信息流范围,该范围联合排序广告和自然内容,并在该范围内选择福利最大化的信息流。然而,存在一个矛盾:策略proofness要求生成的范围与出价无关,即使候选信息流的福利线性依赖于出价。我们的关键技术贡献是一种强化学习方法,它将候选生成和考虑出价的选择纳入训练,使与出价无关的Transformer能够通过考虑单个信息流质量和范围的集体质量来学习生成高福利范围。此外,我们在硬注意力下界定了MIRT类的伪维度,表明近最优期望福利是可学习的,样本复杂度在Transformer大小上呈多项式,在范围大小上仅呈对数。实验上,MIRT在保持精确策略proofness的同时,优于先前非策略proofness的最先进信息流模型。我们的结果表明,基于Transformer的拍卖可以实现考虑外部性的全信息流优化,而不牺牲精确激励兼容性,消除了其实际部署的主要障碍。

英文摘要

Modern online platforms commonly rank ads and organic content separately before blending them into a feed displayed to the user, overlooking externalities: an item's click-through rate depends on its surrounding content, not only on its own position. Recent learning-based feed generation mechanisms model some of these cross-type interactions to globally optimize for the whole feed's welfare. However, these approaches either fix the ordering of organic content, or lack exact strategyproofness guarantees for bidders. To combat these shortfalls, we introduce the Maximal-in-Range Transformer (MIRT) mechanism class, which uses a transformer to generate a range of candidate feeds that jointly order ads and organic content, and selects the welfare-maximizing feed in the range. However, there is a tension: strategyproofness requires the generated range to be bid-independent, even though a candidate feed's welfare depends linearly on the bids. Our key technical contribution is a reinforcement learning approach that incorporates both candidate generation and bid-aware selection into training, enabling a bid-independent transformer to learn to generate high-welfare ranges by accounting for both individual feed quality and the collective quality of the range. Additionally, we bound the pseudo-dimension of the MIRT class under hard attention, showing that near-optimal expected welfare is learnable with sample complexity polynomial in the transformer size and only logarithmic in the range size. Empirically, MIRT outperforms the previous non-strategyproof state-of-the-art feed models while remaining exactly strategyproof. Our results show that transformer-based auctions can deliver externality-aware whole-feed optimization without sacrificing exact incentive compatibility, removing a major obstacle to their practical deployment.

Comments24 pages, 6 figures. Including 10 main pages with 3 figures, and 14 appendix pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑