Noēsis:面向小型本地模型事实关键查询的确定性优先检索与双层上下文水合
Noēsis: Deterministic-First Retrieval with Two-Tier Context Hydration for Factuality-Critical Queries on Small Local Models
浏览论文内容
中文总结 AI 辅助
Noesis提出确定性优先查询平面,通过事实层、位置寻址和双层上下文水合,使20亿参数模型在事实完整性上媲美350亿参数模型,并显著提升检索增强生成的准确性。
中文摘要 AI 辅助
一个错误的数字比没有答案更糟糕。在事实关键领域——媒体中的受众指标、日程安排和版权;医疗保健中的剂量和实验室数值;金融和法律中的数字和引用——一个自信但捏造的值比诚实地承认不确定性更具破坏性。然而,这是我们在小型本地语言模型上观察到的主要失败模式:即使上下文中存在正确的证据,模型也会捏造看似合理的数字和时间戳。最近的工作刻画了这一机制的真实局限:在70亿参数以下,检索增强生成(RAG)的瓶颈不是检索质量而是上下文利用。我们提出Noesis,即Noesis架构的确定性优先查询平面,它在生成之前做出每一个确定性判断。其机制源于摄取架构(另有一项专利申请):(a) 生产者侧事实层,逐字渲染预计算的度量事实而不进行排序;(b) 位置寻址,具有确定性跨源对齐,在查询时间之前以零LLM成本解决;(c) 出处范围界定作为归因约束,具有多级命名引用路由;以及(d) 双层上下文,具有模型触发的逐字水合。在四项消融实验中,一个20亿参数的模型在事实完整性上达到与350亿参数模型相当的水平(所有运行中精确值;在缺失实体陷阱上零捏造数字);结构化检索在20亿参数下比扁平RAG高出11.4个百分点;骨架仅上下文在提示缩小20-30%的情况下保留定量答案;水合在约8秒内恢复逐字叙述,而相比之下约29秒。两个属性对受监管领域很重要:每个查询在单次生成调用中解决,并且每个报告的值在构造上可追溯到其确切来源和位置。
英文摘要
A wrong number is worse than no answer. Across factuality-critical domains -- audience metrics, scheduling and rights in media; dosages and lab values in healthcare; figures and citations in finance and legal -- a confident but fabricated value is more damaging than an honest admission of uncertainty. Yet this is the dominant failure mode we observe on small local language models: even when correct evidence is present in context, models fabricate plausible numbers and timestamps. Recent work characterizes a real limit of this regime: below 7B parameters, the bottleneck of retrieval-augmented generation (RAG) is not retrieval quality but context utilization. We present Noesis, the deterministic-first query plane of the Noesis architecture, which makes every deterministic judgment before generation. Its mechanisms follow from the ingestion architecture (subject of a separate patent application): (a) a producer-side fact layer rendering precomputed metric facts verbatim without ranking; (b) positional addressing with deterministic cross-source alignment, resolved ahead of query time at zero LLM cost; (c) provenance scoping as an attribution constraint with multi-tier named-reference routing; and (d) two-tier context with model-triggered verbatim hydration. Across four ablations, a 2B model reaches parity with a 35B model on factual integrity (exact values in all runs; zero confabulated numbers on absent-entity traps); structured retrieval beats flat RAG by +11.4 points at 2B; skeleton-only context preserves quantitative answers at 20-30% smaller prompts; and hydration recovers verbatim narrative in ~8s versus ~29s. Two properties matter for regulated domains: each query resolves in a single generation call, and every reported value is traceable to its exact source and position by construction.
发表机构
- Alpha Cogs(阿尔法科格斯)
机构由 AI 辅助整理,请以论文原文为准。