arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19831cs.IRcs.AI

复现透明且可审查的推荐:通过自然语言用户画像探索开放权重模型

Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

  • University of Zurich(苏黎世大学)

机构由 AI 辅助整理,请以论文原文为准。

Noah Mamié, Laurin van den Bergh

中文总结 AI 辅助

本研究复现并扩展了基于自然语言用户画像的透明推荐系统研究,验证了UPR的竞争力与透明性,并发现画像扰动仅均匀影响评分而不改变排名,归因于评分回归目标。

中文摘要 AI 辅助

在这项可复现性研究中,我们调查了通过纳入表示用户偏好的生成式自然语言用户画像来增强推荐系统的透明性和可审查性。原始论文探讨了从跨领域(如电影和住宿,即Amazon Movies & TV、TripAdvisor)的原始用户生成评论文本中综合用户画像的方法。至关重要的是,这些自然语言用户画像使得直接的用户交互和干预成为可能,允许用户通过纠正错误归因的偏好或处理冷启动设置来定制推荐。我们成功复现了原始研究的核心发现。此外,我们通过进行系统的上下文消融实验、跨五个不同随机种子的多种子稳定性测试以建立统计可靠性,以及使用nnsight框架进行机制可解释性分析(在反事实画像扰动下探测内部模型表示)来扩展评估。我们的发现验证了原始论文的主张,即用户画像推荐(UPR)在其测试集重排序协议下实现了有竞争力的性能,并使推荐更加透明。扰动自然语言画像确实会改变预测,但它会跨流派均匀地改变预测评分,没有可检测的流派选择性效应,即使在直接激活引导下也保持排名不变。我们将此归因于评分回归目标而非画像接口,而排名目标模型在此任务中明显更优。

英文摘要

In this reproducibility study, we investigate the transparency and scrutability of recommender systems enhanced by incorporating generated natural-language user profiles that represent user preferences. The original paper explores the synthesis of user profiles from raw user-generated review text across domains such as movies and accommodations (Amazon Movies & TV, TripAdvisor). Crucially, these natural-language user profiles enable direct user interaction and intervention, allowing users to customize recommendations by correcting misattributed preferences or addressing cold-start settings. We successfully reproduce the core findings of the original study. Additionally, we extend the evaluation by conducting systematic context ablation experiments, multi-seed stability across five distinct random seeds to establish statistical reliability, and a mechanistic interpretability analysis using the nnsight framework to probe internal model representations under counterfactual profile perturbations. Our findings verify the original paper's claim that User Profile Recommendation (UPR) achieves competitive performance under its test-set reranking protocol and makes recommendations more transparent. Perturbing the natural-language profiles does change predictions, but it shifts predicted ratings uniformly across genres with no detectable genre-selective effect, leaving rankings unchanged even under direct activation steering. We trace this back to the rating-regression objective rather than the profile interface, with ranking-objective models clearly exceeding in this task.

补充信息

↑