arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13785cs.AIcs.LG

分区分数并非系统分数:分解式算法选择中的部署保真度差距

Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection

  • University of Oregon(俄勒冈大学)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiachen Zhang, Yu Tang, Li Zhu

AI总结:

本研究定义部署保真度差距,证明分解式算法选择中分区分数高估可部署系统性能,并在五个基准上验证该差距为正,建议并排报告两类分数。

AI中文摘要:

Oracle式指标,包括虚拟最优求解器、选定组合的VBS、虚拟最优编码以及族内最优摘要,被广泛报告为可部署选择器所能达到的上界。在分解式算法选择中,类似的“分区级分数”允许在所选族内进行Oracle式的最佳算法选择;然而,一旦族选择器被固定,可部署系统必须用学习得到的族内选择器来替代该族内Oracle。我们将部署保真度差距G(R)定义为分区级效用与可部署端到端效用之间的差异,并推导出两个核算后果:一个逐实例的边际遗憾稳定性条件,用于判断分区时的族选择何时是部署最优的;以及一个尖锐的仅分区识别区间,当该区间严格跨越零时,会阻止分区级报告认证可部署的获胜者。在跨越表格型AutoML和组合CSP/SAT的五个公开算法选择基准上,每个分解流水线的G(R)均为正,范围从TabZilla上的0.012到PROTEUS-2014上的0.13。十个分解与扁平决策中有四个的点估计符号发生变化;在PROTEUS-2014上,33点的分区优势缩小为20点的端到端优势。一种训练侧的验证差距校正诊断方法在所有四个符号变化单元上恢复了点估计的可部署符号;它是一种报告辅助工具,而非直接端到端评估的替代品。分区分数和端到端分数应并排报告。

英文摘要:

Oracle-style quantities, including virtual best solvers, selected-portfolio VBS, virtual-best encodings, and best-in-family summaries, are widely reported as upper bounds on what a deployable selector could achieve. In decomposed algorithm selection, an analogous partition-level score grants an oracle choice of the best algorithm within the selected family; once the family selector is fixed, the deployable system must replace that within-family oracle with a learned within-family selector. We define the deployment-fidelity gap G(R) as the difference between partition-level and deployable end-to-end utility and derive two accounting consequences: a per-instance margin-regret stability condition that tells us when a partition-time family choice is deployment-optimal, and a sharp partition-only identification interval that, when it strictly crosses zero, prevents the partition-level report from certifying the deployable winner. Across five public algorithm-selection benchmarks spanning tabular AutoML and combinatorial CSP/SAT, every decomposed pipeline has positive G(R), ranging from 0.012 on TabZilla to 0.13 on PROTEUS-2014. Four of ten decomposed-versus-flat decisions have sign-changing point estimates; on PROTEUS-2014, a 33-point partition advantage shrinks to a 20-point end-to-end advantage. A training-side validation gap-correction diagnostic recovers the point-estimate deployable sign on all four sign-changing cells; it is a reporting aid, not a substitute for direct end-to-end evaluation. Partition and end-to-end scores should be reported side by side.

补充信息

↑