arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型维护的维基知识库的渐进式披露:一项预注册的对比研究

Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation

Theodore O. Cochran

arXiv 2607.04576首次发表:更新:

发表机构

AI for Altruism (A4A)(人工智能促进利他主义组织)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型维护的维基知识库,通过预注册对比研究,改造知识库实现渐进式披露。虽原节省索引加载成本未实现,但答案质量达标且成本降低,源于更有针对性的访问。

AI 中文摘要

大语言模型代理越来越多地根据它们帮助维护的知识库回答问题。一种普遍的直觉认为,渐进式披露(一个紧凑的目录加上每页一行摘要,以便代理只加载它需要的内容)应该比查阅大型整体索引更便宜。我们在一个由大语言模型维护的709页的真实markdown维基上进行了测试。我们对其进行改造以实现渐进式披露,并进行了一项预注册的对比研究,其中语料库的四个版本仅在代理获取内容的方式上有所不同:页面主体在各分支中字节相同,作为不可变的git标签冻结,因此任何测量到的差异都仅归因于访问结构。我们将这些分支与三种访问条件(一个协议受限的代理、一个自由的自路由代理和一个目录预加载机制)交叉,并由一个跨家族评判员根据经过验证的黄金参考对答案进行盲评。一个试点颠覆了前提:一个有能力使用工具的代理从不加载索引,而是从问题中推断页面路径并直接读取,因此改造目标的特定节省并未实现。因此,我们将答案质量作为首要目标,成本作为次要目标。质量不差(检索分支在预注册的范围内与索引基线匹配),而成本在每种机制下都有所下降,从自路由代理的约三分之一到目录预加载下的超过一半,所有置信区间都不包括零。节省并非来自避免索引加载,而是来自更有针对性的访问:检索分支引用的页面更少,工具使用次数更少。该研究同时也是评估有效性的案例研究,将有效性威胁原则应用于产生它的工具。

英文摘要

LLM agents now often answer questions from knowledge bases they help maintain. A common intuition says progressive disclosure should make this cheaper. Instead of loading one large index, the agent reads a compact catalog and one-line page summaries, then opens only the pages it needs. We tested that intuition in a preregistered study on a real 709-page markdown knowledge base maintained by an LLM. We retrofitted it for progressive disclosure and built four versions that differ only in how the agent reaches the pages. The pages themselves are identical in every version, so any difference comes from the access structure alone. Each version was tested three ways, with the agent following a set protocol, choosing its own path, or made to load the catalog first. A judge from a different model family graded the answers blind against verified reference answers. A preparatory pilot changed the question. A capable agent never loaded the large index at all. It worked out from the question where a page was and read it directly. The saving we set out to measure did not exist for such an agent, so we made answer quality the primary outcome. Quality held. Answers from the retrofitted knowledge base were as good as answers from the original, within a margin we set in advance. Two limits apply. Our human rater and the model judge agreed far less than the plan required, so the quality result rests on the judge, backed by sensitivity checks. Quality was also not shown to hold when the agent was forced to load the catalog first, or on the two most reliably graded criteria under a stricter test. Cost fell clearly in every condition we tested, and the retrofitted version cited fewer pages and took fewer tool turns per answer.

Comments15 pages, 3 figures, 6 tables. v2 states its two limits in the abstract. Our human rater and the model judge agreed far less than the plan required. Quality was not shown to hold when the agent had to load the catalog first, or on the two most reliably graded criteria. Preregistered on OSF at https://osf.io/feka7, DOI 10.17605/OSF.IO/FEKA7

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑