发表机构
RMIT University(皇家墨尔本理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大语言模型查询模拟中的知识诅咒问题,提出概念来源框架,可追踪答案侧侵入,实验显示其能以99%的比例消除侵入。
AI 中文摘要
大语言模型生成的搜索查询被广泛用于增强信息检索(IR)评估,但这些查询可能包含预设答案侧文档知识的概念,违反了预搜索用户的信息访问边界。现有验证指标(包括重叠度、多样性和有效性)无法区分罕见的人工定制变体与候选答案侧侵入。我们引入概念来源(concept provenance)这一框架,该框架将查询概念分配给背景支持、人工核心、人工长尾和候选答案侧区域,确立了仅靠检索指标无法检测的边界。将概念来源应用于100个UQV100主题、8个大语言模型(LLM)和5种提示条件下的77004个查询,采用两个提取流程,我们得到跨流程的token级人类中心信息检索(HCIR)斯皮尔曼相关系数ρ为1.0(针对5个条件均值)。候选答案侧概念占非通用概念的7.40%,出现在100个主题中的97个,主题解释了约67%的方差。人工验证得出68.2%的宽松精度,揭示两种机制:知识侵入占45.5%,部署侵入占45.0%。诊断探测显示不成比例的局部检索效应,删除效应量d=-0.47,而随机删除的效应量d=-0.34,但这些概念解释的总评估方差不足2%。因此,概念来源是符合边界的诊断工具,而非评估偏移预测器。在测试条件下,无提示条件能消除侵入;生成后的概念来源选择实现了99%的侵入消除。
英文摘要
LLM-generated search queries are widely used to augment IR evaluation, yet they may contain concepts that presuppose answer-side document knowledge, violating the information-access boundary of pre-search users. Existing validation metrics, including overlap, diversity, and effectiveness, cannot distinguish rare human-tail variation from candidate answer-side intrusion. We introduce concept provenance, a framework that assigns query concepts to backstory-supported, human-central, human-tail, and candidate answer-side zones, operationalizing a boundary that retrieval metrics alone cannot detect. Applying concept provenance to 77,004 queries across 100 UQV100 topics, 8 LLMs, and 5 prompt conditions with two extraction pipelines, we obtain a cross-pipeline token-HCIR Spearman rho of 1.0 over five condition means. Candidate answer-side concepts constitute 7.40 percent of non-generic concepts and appear in 97 of 100 topics, with topic explaining approximately 67 percent of variance. Human validation yields 68.2 percent relaxed precision, revealing two mechanisms: knowledge intrusion at 45.5 percent and deployment intrusion at 45.0 percent. Diagnostic probes show disproportionate localized retrieval effects, with deletion effect size d = -0.47 compared with d = -0.34 for random deletion, but these concepts explain less than 2 percent of aggregate evaluation variance. Concept provenance therefore serves as a boundary-compliance diagnostic rather than an evaluation-shift predictor. Under the tested conditions, no prompt condition eliminates intrusion; post-generation concept-provenance selection achieves 99 percent elimination.
Comments12 pages, 4 figures, and 2 tables. To appear in the Proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26)
Journal refProceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), Rome, Italy, November 7-11, 2026