发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型在与外部知识源结合时的信息辨别问题,提出Learn2Discern框架及基准,通过用户研究和大量模型试验发现模型在来源和真相辨别上存在问题,识别出简单干预措施,发布数据集和调查作为测试平台。
AI 中文摘要
大语言模型越来越多地与互联网等外部知识源一起使用。它们是否能恰当地权衡信息,即对可靠来源更新更多(来源辨别),当断言使先验更接近真相时更新更多(真相辨别)?我们将此形式化为信息辨别,并引入Learn2Discern(L2D),这是一个基于三个规范公理和可解释度量的实验框架及基准。通过预注册、配额匹配的用户研究(n = 299)证实真实的大语言模型用户认可所有三个公理,并报告违反这些公理会降低他们的信任和使用意愿。在13个模型和近670K次试验中,我们发现两个维度上都存在持续失败:模型在来源和真相辨别上表现接近随机,依赖来源受欢迎程度是可靠性的两倍,且无论断言相对于基本事实改善还是恶化其立场,更新大致相同。模型在其先验已经最准确的数据集上最有效地整合外部知识。更新的和更大的模型改善了真相辨别但未改善来源辨别,这是模型复杂性无法解决的盲点。我们识别出能改善两种辨别的简单推理时干预措施。我们发布数据集和调查作为核心对齐属性的测试平台,随着大语言模型取代传统搜索,该属性的重要性不断增加。
英文摘要
LLMs are increasingly used with external knowledge sources like the internet. Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when claims bring priors closer to the truth (truth discernment)? We formalize this as information discernment and introduce Learn2Discern (L2D), an experimental framework and benchmark grounded in three normative axioms with interpretable metrics. To establish external validity, a pre-registered, quota-matched user study (n=299) confirms that real LLM users endorse all three axioms and report that violations reduce their trust and usage intent. Across 13 models and nearly 670K trials, we find consistent failures across both dimensions: models perform near chance on source and truth discernment, rely on source popularity twice as much as source reliability, and update roughly equally whether a claim improves or worsens their position relative to the ground truth. Models integrate external knowledge most effectively on datasets where their priors are already the most accurate. Newer and larger models improve truth discernment but not source discernment, a blind spot that model complexity does not address. We identify simple inference-time interventions that improve both forms of discernment. We release our dataset and survey as a testbed for a core alignment property that scales in importance as LLMs replace traditional search.