发表机构
Tencent LIGHTSPEED(腾讯光速实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出MirageBench,发现12款LLMs普遍存在过度推断用户画像的问题,还揭示了自我监测反转现象,表明外部验证比模型自我报告更可靠。
AI 中文摘要
带有持久记忆的个性化大语言模型(LLMs)正得到越来越多的部署,但其用户模型的忠实度仍未得到检验。我们研究过度推断(OI):即大语言模型伪造超出证据支持范围的用户属性的现象。我们推出MirageBench,它包含150个在刻板印象、反刻板印象和中性画像间平衡的角色,6个跨越“想象梯度”的个性化任务,一个由独立评判者操作的四分类忠实度分类法(已通过400条主张的盲态人工标注者验证:四分类的Cohen's kappa=0.863,二分类的kappa=0.900),以及12个模型、7个系列在143616条经评判的主张上的排行榜。我们发现过度推断普遍存在:12个模型中每一个的主张都有35%至49%存在过度推断(跨模型均值为41.6%;按主张加权为41.8%),本次评估中没有任何模型能避免这一问题。最引人注目的是,我们揭示了自我监测反转:在模型选择层面,模型自我评估的过度推断与评判者测得的过度推断呈负秩相关(rho=-0.60,p=0.044;探索性、宽自举置信区间[-0.90, +0.06],n=12)。自我报告过度推断最少的模型往往被标记为伪造最多,因此自我报告的置信度是比较模型的误导性信号,即便在单一模型内,自我审计仍能对该模型自身的主张进行中等程度的排序(AUROC为0.58至0.83)。我们进一步表明,过度推断具有任务依赖性(27%至59%),且在多轮试点中,推断出的属性以近似线性的方式累积,几乎没有修正。MirageBench主张外部验证而非模型自我报告,作为可信赖个性化的更可靠基础。
英文摘要
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.