上下文层中哪部分在起作用?将语义内容与检索脚手架在Text-to-SQL智能体中的分离
Which Part of the Context Layer Does the Work? Separating Semantic Content from Retrieval Scaffolding in Text-to-SQL Agents
AI总结:
通过四臂消融实验分离语义内容与检索脚手架,证明数据契约中的语义内容主导text-to-SQL准确率提升,优于提示注入知识,并建议语义优先、脚手架其次。
AI中文摘要:
上下文层,即分析智能体在查询时获取的策划文档,在text-to-SQL基准上产生了巨大的准确率提升。有/无对比无法说明该层中哪部分在起作用:是语义内容、传递语义内容的检索脚手架,还是通常伴随的预计算视图。我们在DABStep上对四个模型进行了四臂消融实验,将这三者分离开来。实验工具是一个数据契约:一个携带领域语义和智能体工具所执行规则的YAML工件。其中一臂在冻结契约中清空所有散文字段,同时保持工具表面、检索指令、表允许列表和操作规则逐字节固定不变。将契约自身的SQL表达式编译成视图,给出了预计算层所能达到的上限:在其覆盖的所有176个任务上全部正确。内容占主导地位。它将困难任务准确率从13.9%提升到55.1%,从22.6%提升到56.6%,从22.9%提升到68.4%,从37.0%提升到77.4%,在每个模型上都优于将相同知识粘贴到提示中的做法。没有内容的脚手架在两个flash模型上价值0到5个点,在两个前沿模型上价值14到15个点。与编译后的上限相比,契约臂的不足在于未能进行推导,并且随着模型能力的提升,差距从39个点下降到5个点。契约优于提示,因为它所需的规则只需一次查找即可获得,而不是埋在长提示中:其SQL在高达98%的时间内携带费用语义,而提示臂仅为4%,并且在四个模型中的三个上,每个正确答案的成本最低。对从业者而言:语义优先,脚手架其次,预计算宏仅在智能体确实无法推导时使用。不受约束的臂提交了166条变更语句;受约束的臂则没有。增益仅限于契约的领域。在一个模型上,基准自身保留的金标对契约臂的评分为51.9%,而提示基线为18.8%。
英文摘要:
Context layers, curated documentation that an analytics agent fetches at query time, produce large accuracy gains on text-to-SQL benchmarks. A with/without comparison cannot say which part of the layer does the work: the semantic content, the retrieval scaffolding that delivers it, or the pre-computed views that usually accompany it. We report a four-arm ablation on DABStep on four models that separates the three. The instrument is a data contract: a YAML artifact that carries a domain's semantics and the rules an agent's tools enforce. One arm empties every field of prose in the frozen contract while holding the tool surface, retrieval instruction, table allow-list and operation rules byte-for-byte fixed. Compiling the contract's own SQL expressions into views gives the ceiling a pre-computed layer would reach: gold on all 176 tasks it covers. Content dominates. It raises hard-task accuracy from 13.9% to 55.1%, 22.6% to 56.6%, 22.9% to 68.4% and 37.0% to 77.4%, beating the same knowledge pasted into the prompt on every model. Scaffolding without content is worth 0 to 5 points on two flash models and 14 to 15 on two frontier models. Against the compiled ceiling the contract arm's shortfall is a failure to derive, and it falls from 39 points to 5 with model capability. The contract beats the prompt because the rule it needs is one lookup away rather than buried in a long prompt: its SQL carries the fee semantics up to 98% of the time against the prompt arm's 4%, and at the lowest cost per correct answer on three of four models. For practitioners: semantics first, scaffolding second, pre-computed macros only where an agent demonstrably fails to derive. Ungoverned arms submitted 166 mutating statements; governed arms none. The gain is confined to the contract's domain. On one model the benchmark's own withheld golds grade the contract arm at 51.9% against 18.8% for the prompt baseline.