虚拟集:作为大语言模型生成目标的类型化本体世界,用于有根据的查询和有保障的决策
VirtualSet: Typed Ontology Worlds as an LLM Generation Target for Grounded Queries and Guarded Decisions
浏览论文内容
中文总结 AI 辅助
针对大语言模型操作企业数据时SQL错误信号延迟的问题,提出VirtualSet,通过集合表达式、通用约束投影等技术实现有根据的查询和有保障的决策,在实验中展现出优势,在SQL基准测试中保持竞争力。
中文摘要 AI 辅助
大语言模型越来越多地对企业数据进行读取和操作,但SQL给出错误信号较晚。我们提出了VirtualSet,一种用于大语言模型的实时、接收者类型化的本体世界接口和生成目标。模型发出实体-边世界上的集合表达式,通用约束投影在执行前检查表达式,类型清理读取使用SQL快速路径或有界流解释。相同的基础支持有保障的决策,动作先在模拟世界中运行,世界变化事件在实现前需要外部批准。在BIRD上的实验表明,VirtualSet在有保障决策方面具有优势,且在SQL基准测试中具有竞争力。
英文摘要
Large language models increasingly read and act on enterprise data, but SQL gives a late error signal: hallucinated fields or relations can execute and return plausible wrong answers, while incorrect writes cannot be safely assessed after execution. We present VirtualSet, a live, receiver-typed ontology-world interface and generation target for LLMs. Instead of SQL, the model emits set expressions over entity-edge worlds. Generic Constraint Projection (GCP) checks expressions before execution, while future this preserves concrete receiver types through collection chains, turning invalid fields, edges, receivers, and actions into token-anchored type errors. Type-clean reads use a SQL fast path or bounded stream interpretation, with a parity oracle checking both paths over the exercised operator space. The same substrate supports guarded decisions: actions run first in a simulated world, and world-change events require external approval before actualization. On BIRD, we lift relational schemas into typed worlds and compare VirtualSet with direct SQL while holding the model, evidence, values, zero-shot setting, timeout, glossary, repair/voting, and grader constant where possible. On a frozen 1,072-question split, VirtualSet achieves 67.5% accuracy versus 63.5% for glossary-matched direct SQL with repair and voting (+4.0 points; McNemar exact p = 0.00117) using deepseek-reasoner. Full-corpus analysis finds no engine mis-computation of a type-clean expression; remaining errors arise from model semantics or gold defects. In a 30-body guard corpus, the write chain intercepts 20/20 hallucinated action bodies with zero false positives. VirtualSet thus remains competitive on SQL's home benchmark while providing pre-execution semantics for guarded decisions.