发表机构
Centific Research(森迪菲克研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GVD框架通过双向规则对齐和反事实跨度探测,在本地统一文档版本链接与规则级冲突解决,实现高F1的版本族构建和规则一致性。
AI 中文摘要
文档存储库持续演进。指南和政策会被修订、取代和重新上传,因此相同内容会以不同措辞反复出现,较新的版本会细化或反驳较早的版本。这些不一致属于不断增长的文档集合,而非任何单个文档,然而现有工作将版本管理、重复检测和矛盾检测视为孤立的成对任务,并在一对文档被标记后即停止。我们提出GVD(受控版本管理与去重),这是一个框架,在可审计的更新策略下统一跨文档版本链接与规则级冲突解决。传入文档通过双向规则对齐被分配到版本族,其规则与版本族记忆进行比较,以识别重复、矛盾、不对称细化和新知识,其中反事实跨度探测(CSP)用于解决推理模型误分类为中性关系的相关文档对。关系特定策略抑制重复项,仅将重要变更升级以供审查,同时保留版本谱系作为审计线索。该流程完全在本地运行,无需大型语言模型。在120份企业文档(作为140次摄取,分布于59个版本族)上,GVD在版本族构建上达到0.97的F1分数,在规则级一致性上达到0.94,其中CSP将规则一致性从0.90提升至0.94。
英文摘要
Document repositories evolve continuously. Guidelines and policies are revised, superseded, and re-uploaded, so the same content recurs in different wording and newer versions refine or contradict earlier ones. These inconsistencies belong to the growing collection rather than to any single document, yet existing work treats versioning, duplicate detection, and contradiction detection as isolated pairwise tasks and stops once a pair is labeled. We present GVD (Governed Versioning and Deduplication), a framework that unifies cross-document version linking with rule-level conflict resolution under an auditable update policy. Incoming documents are assigned to version families through bidirectional rule alignment, and their rules are compared against the family memory to identify duplicates, contradictions, asymmetric refinements, and new knowledge, with Counterfactual Span Probing (CSP) resolving related pairs that inference misclassifies as neutral. Relation-specific policies suppress duplicates and escalate only consequential changes for review, retaining version lineage as an audit trail. The pipeline runs fully locally, with no large language model. On 120 enterprise documents processed as 140 ingestions across 59 version families, GVD reaches an F1 of 0.97 for version-family construction and 0.94 for rule-level consistency, with CSP raising rule consistency from 0.90 to 0.94.