arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

个人AI记忆用于评分预测的受控审计

A Controlled Audit of Personal AI Memory for Rating Prediction

Shivam Gupta

arXiv 2610.02764首次发表:更新:

AI 中文总结

本研究通过置换历史评分进行受控审计,区分个人AI在评分预测中依赖关联还是用户倾向,并评估记忆提取与读取器选择的影响。

AI 中文摘要

在结构化评分预测中,个人AI是利用历史物品-评分关联,还是主要依赖用户的评分倾向?我们通过在每个用户内部对历史评分进行置换,同时保持精确的评分分布、物品支持和元数据不变,来审计这一区别。我们将这一控制与完整历史、原生记忆提取以及匹配的数值读取器相结合,在包含Coat和MovieLens数据集的400个留出用户档案和6,160个目标评分的公开冻结评估中进行。在Coat上,经过测试的Qwen编写的Mem0流水线相对于完整历史,将用户宏平均绝对误差分别提高了0.084(Qwen)和0.149(Phi);两个家族调整的bootstrap置信区间均排除零。正确的历史分配对Coat上的两个读取器都有帮助,但相应的MovieLens效应较小且在调整后不具结论性。仅使用历史的岭回归读取器在两个领域均优于Qwen,在MovieLens上优于Phi,而Coat上的Phi比较尚未解决。所有2,800次读取器调用(包括150个无效输出)均在固定的回退规则下保留。一个独立的实现验证了输入、指标以及所有十个主要对比。本研究的贡献是一项可复现的诊断性研究,表明为何提取、关联使用、输出可靠性和读取器选择需要分别评估。

英文摘要

In structured rating prediction, does a personal AI use historical item-rating associations, or mainly the user's rating tendencies? We audit this distinction by permuting historical ratings within each user while preserving the exact rating distribution, item support, and metadata. We combine this control with full history, native memory extraction, and matched numerical readers in a publicly frozen evaluation of 400 held-out user profiles and 6,160 target ratings across Coat and MovieLens. On Coat, the tested Qwen-written Mem0 pipeline increases user-macro mean absolute error relative to full history by 0.084 for Qwen and 0.149 for Phi; both family-adjusted bootstrap intervals exclude zero. Correct historical assignments help both readers on Coat, but the corresponding MovieLens effects are smaller and inconclusive after adjustment. A history-only ridge reader outperforms Qwen in both domains and Phi on MovieLens, while the Coat Phi comparison is unresolved. All 2,800 reader calls, including 150 invalid outputs, are retained under a fixed fallback rule. A separate implementation verifies inputs, metrics, and all ten primary contrasts. The contribution is a reproducible diagnostic study showing why extraction, association use, output reliability, and reader choice require separate evaluation.

Comments15 pages. Code and reproducibility materials: https://github.com/shi1720/personal-ai-memory-studies

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑