摄取时事实编译:面向修订语料库的成本高效且可靠的问答系统
Ingest-Time Fact Compilation for Cost-Efficient and Reliable Question Answering over Revised Corpora
- Endgame Labs, Inc.(Endgame Labs 公司)
- Asia AI Institute(亚洲人工智能研究所)
- Musashino University(武藏野大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出摄取时事实编译架构,在数据摄取时一次性解析修订规则并存储事实记录,使廉价模型在问答时读取编译记录而非重建,显著降低成本并提高可靠性。
AI中文摘要:
大多数智能体问答(QA)系统在最不合适的时间完成其重要的语义工作:每次有人提问时。当语料库包含修订、草稿、撤销、删除以及具有不同权威级别的来源时,模型必须在每次读取时重建受治理的当前状态——然后丢弃该工作并在下一次查询时重复。这有点像数据库在每次有人读取时重建物化视图。我们提出摄取时事实编译,一种在语料库数据被摄取或更改时执行此工作的架构。原始段落被改写为自包含的事实;管理修订、删除、生效日期和来源信任的规则被一次性解析;结果状态被存储为携带来源和修订溯源的类型化记录。在查询时,一个廉价的模型读取编译后的记录,而不是从嘈杂的候选项中重建它。在一个跨五个种子的受控合成实验中,相同的低成本模型在查询时重建下仅在30次试验中的1次产生了正确的值、来源和修订,但在编译底物上的所有30次试验中都成功了,每个问题的平均读取成本降低了12.89倍。在更简单的修订问题上,两种架构都是精确的,但编译路径使用了21.6倍的更少令牌。一个单独的测试发现,事实改写大致将冗长的联邦储备对话减半,同时保持高来源蕴含,但几乎不改变简洁的维基百科散文。这些结果支持一个狭窄但实用的主张:解析一次语料库状态可以使后续的QA对廉价模型更便宜、更可靠。我们发布了开源、MIT许可的实现和实验工件。
英文摘要:
Most agentic question answering (QA) systems do an important part of their semantic work at the worst possible time: every time someone asks a question. When a corpus contains revisions, drafts, revocations, deletions, and sources with different levels of authority, the model must reconstruct the governed current state on every read - then throw that work away and repeat it on the next query. This is a bit like a database that rebuilds a materialized view every time someone reads from it. We present ingest-time fact compilation, an architecture that performs this work when corpus data is ingested or changed. Raw passages are rephrased into self-contained facts; rules governing revisions, deletions, effective dates, and source trust are resolved once; and the resulting state is stored as typed records carrying source and revision provenance. At query time, an inexpensive model reads the compiled record instead of reconstructing it from noisy candidates. In a controlled synthetic experiment across five seeds, the same low-cost model produced the correct value, source, and revision in only one of 30 trials under query-time reconstruction, but in all 30 trials from the compiled substrate, at 12.89 times lower mean read cost per question. On simpler revision questions both architectures were exact, but the compiled path used 21.6 times fewer tokens. A separate test found that fact rephrasing roughly halved verbose Federal Reserve dialogue while preserving high source entailment, but left concise Wikipedia prose essentially unchanged. These results support a narrow but practical claim: resolving a corpus state once can make subsequent QA cheaper and more reliable for inexpensive models. We release the open source, MIT-licensed implementation and experimental artifacts.