发表机构
School of Mathematics and Statistics, Chongqing University; Guangzhou Polytechnic Normal University(重庆大学数学与统计学院; 广州技术师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大型语言模型隐式编码世界模型的风险,提出数据优先本体并构建DaoQL多模态数据库。通过形式化显式世界模型,对比隐式模型优势。实验展示系统性能,如在反事实实验中DaoQL+GPT-4o可组合反事实可分解性大幅提升。
AI 中文摘要
大型语言模型在神经权重中隐式编码世界模型,在医学和金融等高精领域存在幻觉、知识冻结、可解释性差和可修改性差等结构风险。本文提出数据优先本体,将大语言模型视为推理和语言引擎,把确定性知识移入显式多模态数据库DaoQL。形式化了显式世界模型,表明在规则独立性、确定性评估和固定冲突解决下,显式模型为可组合反事实可分解性提供充分条件,隐式模型则缺乏原子读/增量语义。实现的系统聚焦DaoQL的验证存储层和显式评估路径,整合多种引擎。在嵌入式同机设置下给出了相关性能数据,在多域反事实实验中,DaoQL+GPT-4o实现了94%的可组合反事实可分解性,比单独的GPT-4o高49个百分点。本文明确区分了可证明结构、初步经验证据和架构路线图主张。
英文摘要
Large language models encode world models implicitly in neural weights, which exposes four structural risks in high-precision domains such as medicine and finance: hallucination, frozen knowledge, poor explainability, and poor modifiability. This paper proposes data-first ontology: LLMs are treated as reasoning and language engines, while deterministic knowledge is moved into an explicit multimodal database, DaoQL. We formalize an explicit world model and show that, under rule independence, deterministic evaluation, and fixed conflict resolution, explicit models provide a sufficient condition for composable counterfactual decomposability; implicit models lack atomic read/delta semantics and therefore provide no comparable architectural guarantee. The implemented system focuses on DaoQL's verified storage layer and explicit Eval path, integrating graph, column, vector, and full-text engines within one process. KVCache graph nodes, expert hot updates, and the DaoQL-Agent runtime remain future work. On an embedded same-machine setup, DaoQL reports graph BFS at 1.20 ms, HNSW at 83.1 us, and a Fluent hybrid query at 105.8 us; these results indicate engineering potential but must be interpreted with deployment-shape differences from client-server systems. Exploratory measurements on LDBC SNB SF1 and ANN-Benchmarks further show 34/34 query coverage with interactive-class queries mostly in the sub-millisecond to millisecond range, but only 1.8 QPS overall due to long-tail BI/IC queries; ANN-Benchmarks reaches Recall@10 >= 99% at thousand-level QPS after a bridge-edge protection fix. In a five-domain counterfactual experiment (n = 1250), DaoQL+GPT-4o achieves 94% composable counterfactual decomposability, 49 percentage points above GPT-4o alone. The paper explicitly separates provable structure, preliminary empirical evidence, and architectural roadmap claims.
Comments20 pages, 2 figures. Code: https://github.com/zhanbolee/DaoQL-Edu