可审计性原生:面向企业金融领域可信大语言模型分析的本体驱动框架
Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance
浏览论文内容
中文总结 AI 辅助
本文针对企业金融大语言模型应用的信任问题,提出知识驱动分析框架KDAF,结合本体与CARP技术,在FinanceBench评估中,其可审计性优于基线方法,表明可审计性是该框架的核心优势。
中文摘要 AI 辅助
企业在金融领域采用大语言模型的制约因素较少在于流畅度,而更多在于信任:在财务规划与分析(FP&A)及其他受监管的工作流程中,只有当答案可追溯至权威来源且事后可审计时,才具备可用性。本文提出,企业金融领域的检索增强生成模型应在准确性之外,还需以可审计性作为评估标准,并提出知识驱动分析框架(KDAF),该框架通过六个迭代阶段构建本体驱动的知识系统,并通过上下文感知相关性传播(CARP)检索证据,确保每条检索到的事实都携带其关系类型、置信度及来源谱系。在FinanceBench(145个问题)上开展的评估,将KDAF与零上下文推理、BM25、概念加权词汇检索及无依据图遍历进行对比。结果显示:第一,检索是必要的:零上下文推理的正确率为4.1%,而检索增强条件下的正确率为10%-12%;第二,在答案正确性方面,各检索条件的结果无统计学差异(KDAF与BM25的差值为-0.007,95%置信区间为[-0.021, 0.000]),因此仅准确性不足以成为采用结构化检索的理由——本文明确报告了这一负面结果;第三,在可审计性方面,排序发生反转:KDAF的引用可追溯性F1值最高,为0.515,较无依据遍历高出0.027(置信区间[0.006, 0.050]),较BM25高出0.052(置信区间[0.024, 0.083]),区间不包含零。图结构化检索还不会引入问题主题实体之外的证据(426个条目里0个,而词汇基线方法的该比例为16.8%和20.2%),且每个选定条目都能形成完整的溯源链。本文认为,本体驱动检索的成本应基于可审计性而非准确性来衡量。
英文摘要
Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in Financial Planning and Analysis (FP&A) and other regulated workflows, an answer is usable only if it is traceable to authoritative sources and auditable after the fact. This paper argues that retrieval-augmented generation for enterprise finance should be evaluated on auditability alongside accuracy, and presents the Knowledge-Driven Analytics Framework (KDAF), which builds ontology-driven knowledge systems through six iterative stages and retrieves evidence via Context-Aware Relevance Propagation (CARP), so that every retrieved fact carries its relationship type, confidence, and source lineage. An evaluation on FinanceBench (145 questions) compares KDAF against zero-context inference, BM25, concept-weighted lexical retrieval, and ungrounded graph traversal. First, retrieval is necessary: zero-context inference reaches 4.1% correctness against 10-12% for retrieval-augmented conditions. Second, on answer correctness the retrieval conditions are statistically indistinguishable (KDAF vs BM25: -0.007, 95% CI [-0.021, 0.000]), so accuracy alone does not justify structured retrieval here -- a negative result we report explicitly. Third, on auditability the ordering reverses: KDAF attains the highest citation traceability F1 (0.515), exceeding ungrounded traversal by +0.027 (CI [0.006, 0.050]) and BM25 by +0.052 (CI [0.024, 0.083]), intervals excluding zero. Graph-structured retrieval also admits no evidence from outside the question subject entity (0 of 426 items, against 16.8% and 20.2% for lexical baselines), and every selected item resolves to a complete provenance chain. We argue that auditability, not accuracy, is the axis on which ontology-grounded retrieval earns its cost.