arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31890cs.LGcs.AIcs.IR

FARE:面向公平曝光约束与不确定性感知的金融内容个性化深度强化学习

FARE: Deep Reinforcement Learning For Fair Exposure Constrained Uncertainty Aware Financial Content Personalization

Arundeep Chinta, Lucas Vinh Tran, Jay Katukuri

首次发表
浏览论文内容

中文总结 AI 辅助

FARE提出将SOV约束排序建模为深度强化学习,并引入CTR预测不确定性到策略设计,通过模块化执行层实现公平曝光,减少偏差同时最小化参与度损失。

中文摘要 AI 辅助

金融服务中的内容个性化系统必须确保不同产品之间的公平曝光——这一需求源于合同义务,以及防止“富者愈富”动态的必要性,在这种动态中,高点击率(CTR)的内容占据主导地位,而其他相关产品则获得极少的可见性。声量份额(SOV)约束通过保证每个内容类别在顶部位置曝光中占据目标比例,从而促进产品多样性和平衡的用户发现,解决了这一问题。虽然在CTR模型之上使用重排序层在实践中很常见,但我们提出了两个关键创新:(1)将SOV约束排序构建为深度强化学习问题,类似于算法金融中的约束交易执行;(2)明确将CTR预测不确定性纳入智能体的状态空间和策略设计中——使得对于高不确定性预测,可以进行更大的排序调整,因为偏离CTR最优排序的代价较小。我们引入了FARE(公平排序执行器),一个模块化的不确定性感知执行层,它可以将任何黑盒CTR模型的预测转换为SOV公平排序,而无需重新训练底层模型。我们的不确定性加权比例控制策略(FARE-PC)和学习型神经策略(FARE-ES、FARE-PPO)表明,不确定性感知方法可以大幅减少SOV与公平目标的偏差,同时最小化参与度损失,其中无梯度进化策略在合成数据上优于策略梯度方法,而在KuaiRand-Pure上排序则相反。

英文摘要

Content personalization systems in financial services must ensure fair exposure across diverse offerings-a requirement driven by contractual obligations and the need to prevent "rich-get-richer" dynamics where content with high click-through rate (CTR) dominates while other relevant products receive minimal visibility. Share of Voice (SOV) constraints, which guarantee each content category a target fraction of top-position exposure, address this by promoting product diversity and balanced user discovery. While re-ranking layers atop CTR models are common in practice, we propose two key novelties: (1) framing SOV-constrained ranking as a deep reinforcement learning problem analogous to constrained trade execution in algorithmic finance, and (2) explicitly incorporating CTR prediction uncertainty into the agent's state space and policy design-enabling larger ranking adjustments for high-uncertainty predictions where deviation from CTR-optimal ordering is less costly. We introduce FARE (Fair Ranking Executor), a modular uncertainty-aware execution layer that translates any black-box CTR model's predictions into SOV-fair rankings without retraining the underlying model. Our uncertainty-weighted proportional control policy (FARE-PC) and learned neural policies (FARE-ES, FARE-PPO) demonstrate that uncertainty-aware approaches can substantially reduce SOV deviation from fairness targets while minimizing engagement loss, with gradient-free evolution strategies outperforming policy gradient methods on synthetic data and the ordering reversing on KuaiRand-Pure.

补充信息

↑