arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越动作模仿:面向在线广告的决策感知用户模拟器学习

Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising

Zipeng Chen, Jiaer Zheng, Xiangyang Xu, Xinyu Lin, Zhaobin Wang, Zhaohui Liu, Qianjin Xiang, Xiaoyu Zhao, Zhuozhen Yu, Guangshuo Wang, Daxing Chen, Junwei Pan, Zhangbin Zhu, Chengguo Yin, Hao Chen, Tat-Seng Chua, Haijie Gu, Jie Jiang

arXiv 2607.26893首次发表:更新:

AI 中文总结

针对现有用户模拟器仅基于单领域交互模仿动作的缺陷,本文提出决策感知用户模拟器DASH,通过上下文工程、蒸馏思考轨迹与定制规则奖励模型,在腾讯广告数据上验证了其多维度优势。

AI 中文摘要

基于大语言模型(LLM)的用户模拟技术近期取得的进展,已在推荐系统与广告系统的离线评估中展现出应用潜力。然而,现有模拟器通常仅从单领域交互历史中推断用户偏好,且主要以复现点击等可观测动作为优化目标,仅能捕捉用户偏好的部分视角,仅基于动作的预测易引发模型捷径问题,同时会限制模拟的保真度与诊断价值。为应对这些挑战,本文提出DASH——一种决策感知用户模拟器,该模型可从异构跨领域历史中联合生成思考轨迹并预测行为动作。DASH首先引入上下文工程阶段,将异构跨领域历史整合为与决策相关的上下文,并通过提示优化实现对整合后上下文的有效推理;为训练用户模拟器,DASH从强大的LLM中蒸馏思考轨迹作为监督微调(SFT)数据,进一步定制基于规则的奖励模型,该模型从形式、内容、逻辑三个维度评估思考轨迹以用于强化学习(RL)训练,结合动作奖励后,这些信号可共同提升动作预测与思考质量。在覆盖五个异构内容领域的真实腾讯广告数据上开展的大量实验,验证了DASH的有效性、效率、保真度与诊断价值。

英文摘要

Recent advances in LLM-based user simulation have shown promise for offline evaluation of recommendation and advertising systems. However, existing simulators typically infer user preferences from single-domain interaction histories and are primarily optimized to reproduce observable actions such as clicks. Consequently, they capture only a partial view of user preferences, while action-only prediction easily induces model shortcuts and limits both the fidelity and diagnostic value of simulation. To address these challenges, we propose DASH, a decision-aware user simulator that jointly generates thinking traces and predicts behavioral actions from heterogeneous cross-domain histories. DASH first introduces a Context Engineering stage that folds heterogeneous cross-domain histories into decision-relevant context, together with prompt optimization for effective reasoning over the folded context. To train a user simulator, DASH distills thinking trajectories from strong LLMs as SFT data, and further tailors a rubric-based reward model that evaluates thinking traces along form, content, and logic for RL training. Combined with the action reward, these signals jointly improve action prediction and thinking quality. Extensive experiments on real-world Tencent advertising data spanning five heterogeneous content domains demonstrate the effectiveness, efficiency, fidelity, and diagnostic value of DASH.

Comments23 pages, 10 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑