arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SemPlan:面向企业数据的基于大语言模型(LLM)查询的结构化语义规划基准测试

SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data

Bruno Santos Teixeira

arXiv 2608.13612首次发表:更新:

发表机构

Universidade Federal de Ouro Preto (UFOP)(欧鲁普雷图联邦大学(UFOP))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建了SemPlan基准测试,对比四种架构在企业数据LLM查询语义规划中的表现,发现结构化语义规划架构(A3)正确率最高,不同架构在正确率、策略合规性、成本等方面存在权衡。

AI 中文摘要

企业数据的自然语言接口必须将不明确的请求转换为受管控、可执行的行为,同时控制无效查询、策略故障、成本和不确定性。SemPlan基准测试通过一个确定性的合成双语基准评估该架构设计空间,包含1800个英文和巴西葡萄牙语案例,其中1200个案例构成冻结的科学评估子集。在相同模型配置下对比了四种架构:直接SQL生成(A1)、有界工具智能体基线(A2)、结构化语义请求生成后接确定性规划与执行(A3)、澄清/有状态语义规划变体(A4)。在4800条主要记录中,答案正确率绝对值较低:A1为22.25%,A2为22.58%,A3为25.67%,A4为24.25%。A3的观测正确率最高,且在预先指定的配对正确率分析中显著超过A1、A2和A4;A1保持最高的策略正确率和最低的不安全或无效率;A4的平均API成本最低且错误拒绝率最低。在预选的150个案例稳定性子集上,答案正确重复性介于92.00%至98.67%之间。结果支持权衡解读而非通用排名:额外的结构约束改变了故障模式和效率,但未单调提升正确率,也未解决歧义及多轮状态一致性问题。

英文摘要

Natural-language interfaces to enterprise data must translate underspecified requests into governed, executable behavior while controlling invalid queries, policy failures, cost, and nondeterminism. SemPlan Benchmark evaluates this architectural design space with a deterministic synthetic bilingual benchmark containing 1,800 cases in English and Brazilian Portuguese; 1,200 cases form the frozen scientific evaluation subset. Four architectures are compared under the same model configuration: direct SQL generation (A1), a bounded tool-agent baseline (A2), structured semantic-request generation followed by deterministic planning and execution (A3), and a clarification/stateful semantic-plan variant (A4). Across 4,800 primary records, answer correctness was low in absolute terms: 22.25% for A1, 22.58% for A2, 25.67% for A3, and 24.25% for A4. A3 had the highest observed correctness and significantly exceeded A1, A2, and A4 in the pre-specified paired correctness analysis, while A1 retained the highest policy-correct rate and the lowest unsafe-or-invalid rate. A4 had the lowest mean API cost and lowest false-refusal rate. On a preselected 150-case stability subset, answer-correct repeatability ranged from 92.00% to 98.67%. The results support a trade-off interpretation rather than a universal ranking: additional structural constraints changed failure modes and efficiency, but did not monotonically improve correctness or solve ambiguity and multi-turn state consistency.

Comments11 pages, 3 figures, 9 tables. Submitted to Transactions on Machine Learning Research (TMLR)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑