arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WildSEEK:评估语言模型的信息获取能力

WildSEEK: Evaluating Language Models for Information-Seeking

Tanise Ceron, Joachim Baumann, Elisa Bassignana, Berat Cabuk, Dirk Hovy, Debora Nozza

arXiv 2608.30683首次发表:更新:

发表机构

Bocconi University; Stanford University; IT University of Copenhagen; Pioneer Center for AI(博科尼大学; 斯坦福大学; 哥本哈根信息技术大学; 先锋人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有LLM信息获取评估的不足,引入WildSEEK数据集与评估框架,发现超三分之一查询为高风险分析性查询,LLM在多方面存在缺陷,为相关研究提供实证基础。

AI 中文摘要

语言模型正日益成为终端用户获取信息的中介,因此需要对其响应进行系统评估,以构建公平可靠的信息生态系统。然而,现有评估通常针对特定主题或为合成数据,难以捕捉“野外”信息获取查询的复杂性以及模型响应中存在的风险。为解决这一差距,我们引入WildSEEK,这是一个包含3000条来自真实用户交互的信息获取查询的手动标注数据集,以及一个用于评估大语言模型(LLM)生成响应的框架。WildSEEK包含对风险敏感领域(如健康和金融信息)的标注,并区分事实性查询与寻求事实之外答案的分析性查询。我们在WildSEEK上训练分类器,以分析超过180万条真实用户查询。我们发现,超过三分之一的信息获取查询属于高风险,且多为分析性查询。我们的研究结果表明,LLM响应在四个标准上更常失败:谄媚行为、过度依赖、默认以美国为中心的视角以及对弱势群体的处理不当——分析性查询的失败率通常更高。通过提供监测LLM行为的可靠性、安全性和公平性的方法,我们的数据集和评估框架为更广泛的问题提供了实证基础,即当这些系统在信息获取中扮演日益重要的角色时,它们应如何表现。

英文摘要

Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem. Existing evaluations, however, are often topic-specific or synthetic, limiting their ability to capture the complexity of "in the wild" information-seeking queries and the risks present in model responses. To address this gap, we introduce WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses. WildSEEK includes annotations for risk-sensitive domains (e.g. health and financial information), and distinguishes factoid queries from analytical queries which seek responses beyond facts. We train classifiers on WildSEEK to analyze more than 1.8M realistic user queries. We find that over a third of information-seeking queries are high-risk and more often analytical. Our findings show that LLM responses fail more often in four criteria: sycophantic behavior, overreliance, a default US-centric perspective, and poor handling of vulnerable populations -- with failure rates being mostly higher for analytical queries. By providing methods to monitor the reliability, safety, and fairness of LLM behavior, our dataset and evaluation framework offer an empirical foundation for the broader question of how these systems should behave as they take on a growing role in information access.

Comments9 pages, accepted at EMNLP Main 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑