一次对账,随时撰写:用于无漂移、时点研究的信任分层图书馆与多智能体撰写器
Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research
浏览论文内容
中文总结 AI 辅助
该研究提出双层智能体系统,通过信任分层图书馆与多智能体撰写器分离知识库与报告撰写,消除报告矛盾与漂移,实验验证其在多指标、多场景下的有效性与效率优势。
中文摘要 AI 辅助
大型语言模型生成的长篇研究报告存在内容漂移、自相矛盾且失去溯源性的问题:同一指标会出现不同数值,谣言被引用的可信度与经审核的文件相当。本文提出一个双层智能体系统,将维护的时点知识库与报告撰写分离。确定性的“图书馆”将带时间戳的源数据摄入信任分层本体,分层构建证据卡片、权威指标总账和主张图,形成始终最新的真实源,而非基于原始块的每次查询RAG。可移植的多智能体“撰写器”运行时可在任意知识截止时间T生成无矛盾、有证据支撑的报告,仅读取as_of≤T的证据(无超前读取);红队裁决会反馈至图书馆。我们在自行收集的6130个公开源数据语料库上评估,生成555926张证据卡片(涵盖295家发行方和11个行业的SEC EDGAR文件、美国劳工统计局发布内容及维基百科)。从同一库中我们针对不同主题撰写4份时点报告,并开展8项可复现实验,其 headline 指标来自确定性质量控制门,该门通过缺陷注入元评估验证,召回率1.0、精确率1.0。共享指标总账消除6845个横截面矛盾至零。22个黄金案例中,分层优先选择全部正确,而流行度优先基线仅得9/22;信任分层未遗漏任何媒体来源数据,且无政府统计数据取代公司自身文件。红队反驳会反馈并自我修正后续运行,无需人工编辑。在库从235373张卡片增长至555312张的7个截止时间的回放中,无超前读取违规。难度分层模型路由超越全Opus质量上限,运行速度比串行快3.7倍。
英文摘要
Long-form research reports generated by large language models drift, contradict themselves, and lose provenance: the same metric appears with different values, and rumor is quoted as confidently as an audited filing. We present a two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing. A deterministic "librarian" ingests timestamped sources into a trust-tiered ontology, layering evidence cards, an authoritative metric ledger, and a claim graph into an always-current source of truth, not per-query RAG over raw chunks. A portable multi-agent "writer" runtime then composes a contradiction-free, evidence-grounded report at any knowledge cutoff T, reading only evidence with as_of <= T (no look-ahead); red-team verdicts flow back into the librarian. We evaluate on a self-collected, public corpus of 6,130 sources yielding 555,926 evidence cards (SEC EDGAR filings across 295 issuers and 11 sectors, U.S. Bureau of Labor Statistics releases, and Wikipedia). From the one library we compose four point-in-time reports on distinct theses and run eight reproducible experiments, whose headline metrics come from a deterministic quality-control gate, itself validated by defect-injection meta-evaluation at recall 1.0 and precision 1.0. A shared metric ledger removes 6,845 cross-section contradictions to zero. Tier-first selection is correct on 22/22 gold cases where a popularity-first baseline scores only 9/22; trust tiering leaks zero media-sourced numbers, and no government statistic displaces a company's own filing. A red-team refutation propagates back and self-corrects a later run with zero manual edits. Replay exhibits zero look-ahead violations across seven cutoffs while the library grows from 235,373 to 555,312 cards. Difficulty-tiered model routing exceeds the all-Opus quality ceiling while running 3.7x faster than serial.
发表机构
- AWS Generative AI Innovation Center(AWS生成式AI创新中心)
机构由 AI 辅助整理,请以论文原文为准。