arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RPCBench:基于大语言模型的推荐系统中主动前提批判的基准测试

RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation

Zhongru Chen, Yuan Wu, Yi Chang

arXiv 2609.00918首次发表:更新:

发表机构

School of Artificial Intelligence, Jilin University; Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China; International Center of Future Science, Jilin University(吉林大学人工智能学院; 教育部知识驱动人机智能工程研究中心; 吉林大学未来科学国际中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出RPCBench基准测试,用于评估大语言模型的推荐前提批判能力,经对11个大语言模型评估,发现主动检测是该能力的主要瓶颈,且中等长度推理的批判质量最优。

AI 中文摘要

大语言模型正越来越多地被用作交互式推荐助手,因此对它们的评估不应仅停留在合理的物品推荐层面,还需测试其能否识别有缺陷的推荐请求。现有推荐基准主要评估排序、生成或偏好满足情况,而现有的错误检测基准通常不基于推荐特有的用户和候选证据。为解决这一差距,我们推出RPCBench,一个用于评估推荐前提批判(Recommender-Premise Critique)的基准测试:即检测、诊断并妥善处理自然语言推荐请求中错误前提的能力。RPCBench包含来自五个推荐领域的基于证据的测试实例,涵盖十种前提错误类型。每个实例提供可见的推荐上下文和损坏的用户查询。我们还设计了一个细粒度评估框架,用于衡量主动检测、错误定位、检测后处理策略以及证据忠实度。通过对11个大语言模型的系统评估,我们发现主动检测是推荐前提批判的主要瓶颈,且模型在未明确说明的前提错误上表现最差。我们还观察到,目标关键信息密度比冗余证据更重要,且更长的推理并不单调提升批判质量:性能在中等推理长度时达到峰值,而过长的推理会伴随过度思考惩罚。代码可在该https URL获取。

英文摘要

Large language models are increasingly used as interactive recommender assistants. Their evaluation should therefore go beyond plausible item recommendation and test whether they can recognize flawed recommendation requests. Existing recommender benchmarks mainly assess ranking, generation, or preference satisfaction, while existing error-detection benchmarks are usually not grounded in recommendation-specific user and candidate evidence. To address this gap, we introduce RPCBench, a benchmark for evaluating Recommender-Premise Critique: the ability to detect, diagnose, and properly handle faulty premises in natural-language recommendation requests. RPCBench contains evidence-grounded test instances from five recommendation domains and covers ten types of premise failures. Each instance provides a visible recommendation context and a corrupted user query. We further design a fine-grained evaluation framework that measures proactive detection, error localization, post-detection handling strategy, and evidence faithfulness. Through a systematic evaluation of 11 LLMs, we find that proactive detection is the main bottleneck in Recommender-Premise Critique, and models perform worst on underspecified-premise errors. We also observe that target-critical information density matters more than redundant evidence, and that longer reasoning does not monotonically improve critique quality: performance peaks at intermediate reasoning length, while overly long reasoning is accompanied by an overthinking penalty. The code is available at https://github.com/ZhongruChen/RPCBench.

Comments45 pages, 8 figures. Code available at https://github.com/ZhongruChen/RPCBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑