从提示工程到行为对齐:用于推荐评估的个性化大语言模型评判者
From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation
- Netflix(网飞公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对离线推荐评估中LLM存在的双向合理化问题,提出序列行为对齐框架,经真实日志验证,其Macro-F1较零样本基线提升32.19%,可缓解失效模式且无需人工管道开销。
AI中文摘要:
传统离线推荐评估严重依赖复杂且需人工维护的特征管道,难以扩展。大语言模型(LLM)提供了直接从原始文本日志预测用户参与度的可行替代方案,但本研究的实证分析发现了一种称为双向合理化的关键失效模式:在零样本设置下,LLM会用完全相同的证据,分别为同一物品的正面和负面用户参与结果进行有说服力的论证,凸显了现成LLM在预测用户参与度时的不可靠性。为解决该问题,我们开发并应用了序列行为对齐框架,将微调与配对正确及反事实理由的偏好优化相结合。在真实主页交互日志上的评估显示,该对齐推理方法的Macro-F1分数较零样本基线提升了32.19%,且与基于特征工程的生产基线表现相当。结果表明,行为对齐可缓解双向合理化问题,同时提供人类可解释的推理轨迹,且无需人工管道开销。
英文摘要:
Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact same item with identical evidence, highlighting the unreliability of off-the-shelf LLMs in predicting user engagement. To resolve this, we develop and apply a sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales. Evaluated on real-world homepage interaction logs, this aligned reasoning approach achieves a 32.19\% lift in Macro-F1 score over the zero-shot baseline and matches the production feature-engineered baseline. The results demonstrate that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pipeline overhead.