发表机构
The Chinese University of Hong Kong; Xi’an University of Electronic Science and Technology(香港中文大学; 西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对跨模型计算复用中的质量不确定性和耦合调度难题,提出QARS在线决策算法,联合优化结果准备与任务分配,实现遗憾上界并降低总成本。实验显示成本降低18.0%,遗憾降低63.9%。
AI 中文摘要
一个模型计算的中间结果可以被其他模型复用以执行其任务。现有工作主要关注实际执行,在优化复用决策方面留下了理论空白。这一优化面临两个挑战:质量不确定性,因为复用对任务质量的影响在不同模型间是不确定的;以及耦合调度,因为任务需要分担准备可复用结果的成本。这些挑战相互叠加:质量必须在线学习,但共享的准备结构即使已知质量也使调度成为NP难问题,打破了现有方法中的关键假设。我们将跨模型计算复用形式化为一个在线决策问题,并开发了质量感知复用调度(QARS)算法来解决它。对于质量不确定性,QARS从选定的、可能延迟的反馈中学习任务相关的复用质量,并使用乐观估计来指导决策。对于耦合调度,它联合选择准备哪些结果以及哪些任务应使用它们,使调度精度适应剩余的质量不确定性。对于所考虑的问题,我们的分析在遗憾中分离了学习和优化误差,并量化了调度精度与计算之间的权衡。完成质量感知停止规则可产生$\widetilde O(\sqrt{T})$的遗憾,同时保持可行性。实验证明了QARS在优化跨模型复用方面的有效性,将计算和质量损失的总成本降低了最多18.0%,平均遗憾比最强的调度基线降低了63.9%。
英文摘要
An intermediate result computed by one model can be reused by other models to perform their tasks. Existing work mainly focuses on practical execution, leaving a theoretical gap in optimizing reuse decisions. This optimization faces two challenges: quality uncertainty, because the effect of reuse on task quality is uncertain across models, and coupled scheduling, because tasks need to share the cost of preparing reusable results. These challenges compound each other: quality must be learned online, but the shared preparation structure makes scheduling NP-hard even with known quality, breaking the key assumption in existing methods. We formulate cross-model computation reuse as an online decision problem and develop the Quality-Aware Reuse Scheduling (QARS) algorithm to address it. For quality uncertainty, QARS learns task-dependent reuse quality from selected, possibly delayed feedback and uses optimistic estimates to guide decisions. For coupled scheduling, it jointly chooses which results to prepare and which tasks should use them, adapting scheduling accuracy to the remaining quality uncertainty. For the considered problem, our analysis separates learning and optimization error in regret and quantifies the tradeoff between scheduling accuracy and computation. Completing the quality-aware stopping rule yields $\widetilde O(\sqrt{T})$ regret while preserving feasibility. Experiments demonstrate the effectiveness of QARS in optimizing cross-model reuse, reducing the combined cost of computation and quality loss by up to 18.0%, and mean regret by 63.9% over the strongest scheduling baseline.