arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CATune:面向数据库管理系统配置调优的结构约束感知贝叶斯优化

CATune: Structural Constraint-Aware Bayesian Optimization for DBMS Configuration Tuning

Fangping Lan, Qi Zhang, Eduard Dragut

arXiv 2610.09276首次发表:更新:

发表机构

Temple University(天普大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CATune提出一种结构约束感知的贝叶斯优化框架,显式建模DBMS配置旋钮间的确定性排序约束,在约束一致子空间内优化,显著提升调优效率与质量,实验显示吞吐量最高提升63.37%。

AI 中文摘要

现代数据库管理系统(DBMS)暴露了数百个配置旋钮,导致搜索空间高维且异构,使得自动调优成本高昂。现有的基于机器学习的调优系统通常将配置域视为盒约束,并依赖工作负载反馈隐式捕获旋钮间的关系。然而,数据库管理系统文档指定了确定性的旋钮依赖约束,特别是排序约束,这些约束刻画了配置空间中结构有效的区域。我们提出了CATune,一个约束感知的贝叶斯优化(BO)框架,将确定性的旋钮间排序约束建模为搜索域的结构组成部分。CATune不是在采样违规中学习可行性边界,而是在约束一致的子空间内进行优化。我们开发了一种拓扑感知的采样策略,在探索过程中尊重依赖结构,并避免事后约束处理的低效性。为了实现自动约束发现,我们进一步设计了一个精度优先的提取管道,结合基于LLM的解析与可靠性保障,以减轻幻觉依赖。在PostgreSQL和MySQL上使用TPC-C和TPC-H工作负载的实验表明,CATune在代理模型和BO框架中显著提高了样本效率和最终调优质量。在默认范围内,CATune达到基线最优值的速度最高提升12.5倍,吞吐量最高提升63.37%。在知识引导的缩减范围和替代优化实现下,改进仍然持续。这些结果表明,显式建模系统定义的确定性排序约束增强了优化鲁棒性和系统稳定性。

英文摘要

Modern DBMSs expose hundreds of configuration knobs, resulting in a high-dimensional and heterogeneous search space that makes automated tuning costly. Existing ML-based tuning systems typically treat the configuration domain as box-constrained and rely on workload feedback to implicitly capture inter-knob relationships. However, DBMS documentation specifies deterministic knob dependency constraints, particularly ordering constraints, that characterize structurally valid regions of the configuration space. We present CATune, a constraint-aware Bayesian optimization (BO) framework that models deterministic inter-knob ordering constraints as structural components of the search domain. Instead of learning feasibility boundaries through sampled violations, CATune performs optimization within a constraint-consistent subspace. We develop a topology-aware sampling strategy that respects dependency structure during exploration and avoids the inefficiencies of post-hoc constraint handling. To enable automated constraint discovery, we further design a precision-first extraction pipeline that combines LLM-based parsing with reliability safeguards to mitigate hallucinated dependencies. Experiments on PostgreSQL and MySQL using TPC-C and TPC-H workloads show that CATune substantially improves both sample efficiency and final tuning quality across surrogate models and BO frameworks. Under default ranges, CATune reaches the baseline optimum up to 12.5x faster and improves throughput by up to 63.37%. The improvements persist under knowledge-guided reduced ranges and alternative optimization implementations. These results demonstrate that explicitly modeling system-defined deterministic ordering constraints enhances optimization robustness and system stability.

Comments14 pages including references, 10 figure, 2 tables. Accpeted by PVLDB, Volume 19, 2026

DOI:10.14778/3849398.3849421

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑