发表机构
University of California San Diego; Meta AI(加利福尼亚大学圣迭戈分校; Meta AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出个人代理中介推荐范式,通过PAMO优化方法利用跨平台历史平衡平台修正与有害覆盖,在MediateRec基准上验证了有效性。
AI 中文摘要
现代推荐正在从以平台为中心的个性化转向用户管理的个性化,其中个人LLM代理可以代表用户跨服务行事。我们将这一新兴范式形式化为个人代理中介推荐:平台推荐器使用平台本地信息对候选集进行排序,个人代理使用用户授权的跨平台历史来中介生成的排序并产生最终的Top-K列表。这种中介并非易事:平台排序可以编码个人代理无法观察到的强群体证据,因此有效的中介必须在有益的补救和有害的覆盖之间取得平衡。为了研究这种权衡,我们引入了MediateRec,一个基准,包括可扩展的代理跨平台环境和在受控平台-代理信息边界下的真实跨平台测试。为了训练代理有效使用跨平台历史,我们进一步提出了个人归因中介优化(PAMO),它反事实地掩盖该历史以估计个人中介支持,并在平台相对价值下限下重新分配排名感知的优势质量。我们从理论上证明PAMO保留了截止级别的优势质量,并且在保持该质量而不降低平均平台相对价值的一阶重新分配中是局部最优的。在MediateRec上的实验表明,个人代理中介能够实现有意义的平台修正,但即使是强大的专有LLM也会引入不可忽视的有害覆盖。PAMO在可见和不可见的目标平台以及真实跨平台测试中始终优于匹配的结果仅RL,同时实现了更好的补救-危害平衡。
英文摘要
Modern recommendation is shifting from platform-centric personalization toward user-governed personalization, where a personal LLM agent can act on the user's behalf across services. We formalize this emerging paradigm as Personal-Agent Mediated Recommendation: a platform recommender ranks a candidate set using platform-local information, and a personal agent uses user-authorized cross-platform history to mediate the resulting ranking and produce the final top-K slate. Such mediation is nontrivial: the platform ranking can encode strong population evidence that the personal agent cannot observe, so effective mediation must therefore balance beneficial rescues against harmful overrides. To study this trade-off, we introduce MediateRec, a benchmark that includes scalable proxy cross-platform environments and a real cross-platform test under a controlled platform-agent information boundary. To train the agent to use cross-platform history effectively, we further propose Personal Attribution Mediation Optimization (PAMO), which counterfactually masks that history to estimate personal mediation support and reallocates rank-aware advantage mass under a platform-relative value floor. We theoretically prove that PAMO preserves cutoff-level advantage mass and is locally optimal among first-order reallocations that preserve this mass without lowering average platform-relative value. Experiments on MediateRec show that personal-agent mediation enables meaningful platform corrections, yet even strong proprietary LLMs introduce non-negligible harmful overrides. PAMO consistently improves over matched outcome-only RL across seen and unseen target platforms and on the real cross-platform test, while achieving a better rescue-harm balance.