发表机构
University of Nebraska–Lincoln(内布拉斯加大学林肯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出监督式结构学习方法训练知识库,以智能体策划文档存储,提升动作效率与准确率,其泛化性与键覆盖相关,欠训练的存储可通过增加训练问题进一步优化。
AI 中文摘要
检索增强生成将文档存储视为固定输入,而让智能体策划文档存储的系统从未衡量过策划对存储的作用。我们反转了框架:知识库即模型。训练智能体针对当前存储回答一个监督式问题,查看标准答案后编辑存储;后续在固定动作预算下对冻结快照上的未改动阅读器进行测试。离线图构建是无监督的,而(问题,答案)对是我们的标签——正是这种监督让结构变得廉价。就每索引的语料库条目而言,它相比覆盖所有内容的无监督实体索引,动作节省量是后者的1.6倍,准确率是后者的1.8倍,使用的链接数为1913个,而后者为196112个。在训练过的存储的问题上,未改动阅读器的动作减少了31%,准确率更高,且该结果在官方PhantomWiki生成任务上可复现,我们并未编写其问题。为衡量其覆盖范围,我们引入了键覆盖梯度,这是一种探测方法,可改变训练集触及问题的程度,将训练/测试拆分的通过/失败替换为衰减曲线。泛化证明依赖于端点:准确率可迁移到未见过的问题(当问题的两个键都被索引时,F1提升0.167;一个键被索引时,提升0.100;无键被索引时,无提升),而动作节省仅保留在训练过的问题上。由于衰减由覆盖范围而非新颖性索引,更多训练可扩展其效果——且存储处于欠训练状态,未饱和:覆盖范围随新问题线性增长,且在训练重复时停止,因此100个问题可覆盖语料库的四分之一,若问题数增加四倍则可缩小差距。
英文摘要
Retrieval-augmented generation treats the document store as a frozen input, and the offline pipelines that do build structure over it build it unsupervised -- a whole corpus indexed at uniform effort, with no signal about which structure a question will need. We instead treat the knowledge base as a non-parametric model trained on (question, answer) pairs: a curator agent answers a supervised question against the current store, is shown the gold answer, then edits the store. The store carries forward, and we evaluate the curated store with a test set, on two contamination-free benchmarks: KBGym, a fictional-universe generator we release, and PhantomWiki. Generalization is probed with four question groups of decreasing overlap with the training set: the trained questions themselves, and unseen questions sharing both of their keys with training, one key, or neither. The curated store's advantage grows with overlap -- from parity where no key was shared, through +0.176 F1 where both keys were, to 25% fewer actions at +0.294 F1 on the trained questions, the one cell significant on both benchmarks -- while matching HippoRAG's gains with 1,913 links against its 196,112: per point of corpus covered, 1.5x the action saving and 2.1x the accuracy gain. Accuracy rises steadily with the share of the corpus the indexes cover, so training on more questions widens coverage, and with it the generalization.
Comments10 pages, 4 figures, 5 tables. Submitted to IEEE BigData 2026