arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11493cs.AIcs.MA

从文档孤岛到流程智能:用于CMC工艺开发的多层知识图谱

From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development

发表机构赛诺菲美国
查看机构详情
  • Sanofi US(赛诺菲美国)

机构由 AI 辅助整理,请以论文原文为准。

Reza Amirmoshiri, Faryad Sahneh, Yasser Jangjou

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种模块化智能体AI平台,通过构建双层知识图谱(词汇层与智能层)整合CMC工艺开发文档,解决知识碎片化问题,并在505个问题上验证了高准确率与可靠性。

中文摘要 AI 辅助

化学、制造和控制(CMC)工艺开发在从药物发现到商业生产的多阶段、知识密集型连续过程中产生了大量的技术信息。这些知识传统上分散在不同职能部门和异构格式中,导致在技术转移和监管申报期间出现可追溯性缺口和显著的知识管理成本。我们提出一个模块化的智能体AI平台,可将异构的工艺开发文档语料库转换为可查询的双层知识图谱。基础知识层通过无损摄取数字、扫描、手写和多语言文档,构建具有文档-章节-块(Document-Section-Chunk)层次结构的词汇图;智能层则提取与本体对齐的实体,并通过基于溯源锚定的领域图桥接跨文档概念。LLM智能体在这两层上运行,为每个问题选择最合适的检索路径。我们使用一种新颖的三层协议评估词汇层,该协议衡量检索增强生成(RAG)系统在专有数据上的部署保真度,并在来自赛诺菲小分子项目的38份开发报告中精选的505个问题上进行了演示。第一层多项选择准确率为95%,表明平台可靠性强;更严格的第二层LLM评判通过率为85%,该指标在比较性和全语料库问题上有所下降,揭示了一个仅靠第一层准确率无法捕获的失败分类。一个路由智能体根据问题类型在层之间进行选择。我们预计该协议将使未来的智能体平台设计者能够针对非公开数据库评估其系统,并且基于图的架构将在制药行业得到更广泛的采用,作为将分散的文档存储库转变为结构化流程智能的一种手段。

英文摘要

Chemistry, Manufacturing and Controls (CMC) process development generates an enormous body of technical information across a multi-stage, knowledge-intensive continuum from drug discovery to commercial manufacturing. This knowledge is traditionally fragmented across functions and heterogeneous formats, causing traceability gaps and significant knowledge-management costs during technology transfer and regulatory filing. We present a modular agentic-AI platform that converts a heterogeneous corpus of process-development documents into a queryable, dual-layer knowledge graph. A base knowledge layer builds a lexical graph with a Document-Section-Chunk hierarchy through lossless ingestion of digital, scanned, handwritten, and multilingual documents, while an intelligence layer extracts ontology-aligned entities and bridges cross-document concepts through a provenance-anchored domain graph. LLM agents operate across both layers, selecting the retrieval path best suited to each question. We evaluate the lexical layer with a novel three-tier protocol measuring the deployment-fidelity of a retrieval-augmented generation (RAG) system on proprietary data, demonstrated on 505 questions curated from 38 development reports of a Sanofi small-molecule program. Tier-1 multiple-choice accuracy of 95% signals strong platform reliability; the stricter Tier-2 LLM-judge pass rate of 85%, which degrades on comparative and corpus-wide questions, reveals a failure taxonomy that Tier-1 accuracy alone fails to capture. A router agent selects between layers according to question type. We anticipate this protocol will enable future designers of agentic platforms to assess their systems against nonpublic databases, and that graph-based architectures will see broader adoption in pharma as a means of transforming fragmented document repositories into structured process intelligence.

↑