arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DBA-Bench:基于大语言模型的数据库操作代理的生产保真度基准测试

DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng

arXiv 2607.22165首次发表:更新:

发表机构

University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于大语言模型的数据库操作代理评估问题,提出DBA-Bench基准测试,通过生产保真度等解决评估与生产操作差距,含多场景和难度标签,评估多基线组,揭示安全端到端修复困难。

AI 中文摘要

基于大语言模型的数据库代理虽展现出潜力,但不同的任务范围、测试平台和指标阻碍了比较。我们识别出评估与生产操作之间的四个差距:实时环境保真度、观测空间规模与复杂性、解决方案空间开放性、场景复杂性与覆盖范围。我们提出了DBA-Bench基准测试,通过生产保真度、结果优先评估和可控场景可重复性来解决这些差距。它使用带有活跃工作负载、持久状态和多源观测的仪器化PostgreSQL环境;在安全约束下通过可测量的恢复或故障消除来定义成功;每次运行前通过特定场景检查恢复快照。该基准测试包含七个任务领域的106个场景,基于参考路径诊断深度和环境复杂性有两个公开难度标签。我们评估了九个基线组,包括六个基础模型系统、两个由GPT-5.5支持的数据库代理和一个人类数据库管理员参考。在848次自动运行中,诊断率、成功率和安全通过率分别为32.7%、19.6%和12.4%;最佳自动基线的安全通过率为17.9%,而人类数据库管理员参考为93.4%。自动安全通过率从简单场景的19.6%降至困难场景的7.6%,凸显了安全端到端修复的难度。

英文摘要

LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a running database); observation-space scale and complexity (causal diagnosis across thousands of time series, business logs, and concurrent activity); solution-space openness (multiple remediations with different operational trade-offs); and scenario complexity and coverage (faults cascading across internal mechanisms and operational domains). We present DBA-Bench, a benchmark addressing these gaps through production fidelity, outcome-first evaluation, and controlled scenario reproducibility. It uses instrumented PostgreSQL environments with active workloads, persistent state, and multi-source observations; defines success by measurable recovery or fault elimination under safety constraints; and restores snapshots with scenario-specific checks before each run. The benchmark contains 106 scenarios across seven task domains, with two public difficulty labels based on reference-path diagnostic depth and environmental complexity. We evaluate nine baseline groups, including six foundation-model systems, two GPT-5.5-backed database agents, and a Human DBA reference. Across 848 automated runs, Diagnosis, Outcome, and Safe Pass rates are 32.7%, 19.6%, and 12.4%; the best automated baseline reaches 17.9% Safe Pass versus 93.4% for the Human DBA reference. Automated Safe Pass falls from 19.6% on Easy scenarios to 7.6% on Hard scenarios, underscoring the difficulty of safe end-to-end remediation.

Comments14 pages, 6 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑