AI 中文总结
EvoWiki是一种增量问答架构,通过解耦构建与读取、采用状态覆盖协议等技术,在跨会议知识演化任务中提升了事实一致性与可追溯性,在多数据集上准确率显著优于基线。
AI 中文摘要
在跨越多次会议的长期协作中,决策、风险等事实状态会不断被修订、推翻和替换。现有的长上下文方法通常会堆叠全部历史内容,而许多检索增强生成(RAG)和结构化记忆方法则将知识组织为静态或仅追加的事实,并在读取时依赖语义相关性。由于未对知识生命周期进行显式建模,这些方法可能会同时保留新旧冲突的状态,或丢弃历史内容,进而导致检索过时、答案难以验证。我们提出EvoWiki,一种面向动态长文本的增量问答架构,将离线增量构建(BUILD)与在线结构化读取(READ)解耦。BUILD捕捉从提案到决策的会议内微观演化,采用实体版本链与细粒度状态覆盖协议,明确区分当前有效状态与已被取代的历史,同时保留会议级溯源锚点。READ绕过基于相关性的Top-k检索,在完整Wiki上执行确定性实体寻址、时间解析与跨实体多跳聚合,以生成有依据且可追溯的答案。我们还推出CrossMeet,一个高保真双语基准,旨在模拟长期状态演化,涵盖事实一致性、时间推理与跨会议多跳推理。在六个数据集和两个读取器模型上,EvoWiki的宏观平均判断准确率较最强基线分别提升9.72和10.00个百分点。人工评估显示,在频繁状态翻转场景下,EvoWiki更鲁棒且事实忠实度更高,验证了面向有效状态的读取是应对跨会议知识演化的可靠方法。
英文摘要
In long-term collaboration spanning multiple meetings, factual states such as decisions and risks are continually revised, overturned, and replaced. Existing long-context methods typically stack the entire history, while many RAG and structured-memory methods organize knowledge as static or append-only facts and rely on semantic relevance at read time. Without explicit modeling of knowledge lifecycles, these approaches may retain conflicting old and new states simultaneously or discard history, leading to stale retrieval and answers that are difficult to verify. We present EvoWiki, an incremental question-answering architecture for dynamic long-form text. EvoWiki decouples offline incremental construction (BUILD) from online structured reading (READ). BUILD captures the intra-meeting micro-evolution from proposal to decision and uses entity version chains and a fine-grained State-Overwrite Protocol to explicitly distinguish current valid states from superseded history while preserving meeting-level provenance anchors. READ bypasses relevance-based Top-k retrieval and performs deterministic entity addressing, temporal resolution, and cross-entity multi-hop aggregation over the complete Wiki to produce grounded and traceable answers. We further introduce CrossMeet, a high-fidelity bilingual benchmark designed to simulate long-term state evolution, covering factual consistency, temporal reasoning, and cross-meeting multi-hop reasoning. Across six datasets and two reader models, EvoWiki improves macro-average Judge Accuracy over the strongest baselines by 9.72 and 10.00 percentage points, respectively. Human evaluation shows that EvoWiki is more robust and factually faithful under frequent state flips, validating valid-state-oriented reading as a reliable approach to cross-meeting knowledge evolution.
Comments12 pages, 5 figures, 10 tables