发表机构
Adobe Media & Data Science Research(Adobe媒体与数据科学研究)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大型语言模型个性化能力,将贝叶斯说服框架应用于生成式智能体并实例化到销售场景,发布SDR-Bench语料库,发现个性化存在平台期,通过现场部署验证框架,还公开了SDR-Arena和SDR-Bench以支持相关可重复研究。
AI 中文摘要
个性化是在保持发送者、渠道和时间固定的情况下,改变信息以促使特定接收者采取行动的行为,在心理学和营销中由来已久。大型语言模型通过根据推断的接收者状态生成连续的消息变体,消除了经典检索和排序方法的有限库存约束。现有基准测试衡量的是发送者端的适应,而第三方是否被诱导采取行动的两方问题仅通过A/B测试和小规模人类研究进行过调查。本文将贝叶斯说服框架应用于生成式智能体,并在销售场景中实例化。发布了SDR-Bench语料库,观察到个性化存在平台期,在财富100强科技队列中没有模型能在统计上区分成功与不成功的推广。通过与12名专业销售代表的现场部署验证了框架,发布SDR-Arena和SDR-Bench以支持大规模生成式个性化的可重复研究。
英文摘要
Personalization is classically a two-party problem: a sender chooses what to say, and a receiver with independent objectives decides whether to act. A salesperson pitching the same analytics product leads with HIPAA compliance for a hospital and real-time reporting for a retailer, expecting a different argument to work on each. Existing LLM personalization benchmarks measure a narrower, one-party property: whether output matches the preferences of the same user it serves-sender and receiver being the same, as when RLHF aligns an assistant to its own user. The two-party case is harder to study automatically, since it needs ground truth linking specific content to an observed receiver action. Sales outreach provides this: a message written for one prospect, recorded against whether it produced a reply, a call, or a closed deal. We introduce SDR-Arena, a framework for benchmarking two-party generative personalization at scale, and SDR-Bench, a public corpus of 50,000 customer success stories across 22 industries and 3500 enterprises. Given only pre-outcome information, an agent must reconstruct the arguments that won the deal, scored by a weighted nugget-recall metric (WCS). The best model, Claude Sonnet 4.6, reaches 55.8% WCS indicating it recovers only half the winning content-a plateau we observe across model families that costly deep-research pipelines do not close. An ablation shows the cause is retrieval, not reasoning: models improve substantially given the facts a human researcher would gather, but rarely find them through web search alone. Two studies with professional SDRs support the metric: only 48% of generated pitches were rated usable without editing, and WCS yields model rankings consistent with evaluation against expert strategies authored independently of any model output.