发表机构
IIT Delhi; University of Chicago; University of Missouri; IISc Bangalore(印度理工学院德里分校; 芝加哥大学; 密苏里大学; 印度科学研究院班加罗尔)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlexiFlow提出基于多臂老虎机的动态模型切换方法,根据运行时间和断言通过率排序模型,在运行时切换并重用中间结果,提升ML工作流准确率高达23%和效率48%。
AI 中文摘要
模型优化有助于提升机器学习工作流的推理性能和准确性。然而,依赖单一模型对所有数据批次进行推理,往往无法最大化准确性,从而影响整体性能。在许多情况下,替代模型可能在主模型表现不佳的特定数据子集上表现更好。我们在真实机器学习工作流上的实验确实表明,切换模型可将工作流准确性提升高达23%。然而,现有系统缺乏基于性能自适应切换模型的能力,迫使用户手动按顺序测试模型。我们提出了FlexiFlow,一个数据流系统,当当前模型表现出低准确性时,动态地在替代模型之间切换。FlexiFlow使用一种新颖的多臂老虎机方法学习模型排序,该方法考虑了模型运行时间、通过用户定义断言的概率以及机器学习工作流的计算结构。我们表明,标准的汤普森采样方法不足以在机器学习工作流中切换模型。相比之下,我们提出的方法有效且可扩展至复杂的真实世界机器学习工作流。实验表明,在运行时切换模型并重用中间结果,不仅提供了更高的准确性,而且与顺序工作流运行相比,还实现了48%的效率提升。
英文摘要
Model optimizations help improve inference performance and accuracy of ML workflows. However, relying on a single model to perform inference across all data batches often fails to maximize accuracy and thus overall performance. In many cases, alternate models could perform better on specific subsets of data where a primary model underperforms. Our experiments with real ML workflows indeed show that switching models improves workflow accuracy by up to 23%. Yet, current systems lack the ability to adaptively switch between models based on performance, forcing users to manually test models in sequence. We present FlexiFlow, a dataflow system that dynamically switches between alternate models when the current model exhibits low accuracy. FlexiFlow learns to rank models using a novel multi-armed bandit approach that accounts for model runtimes, probability of passing user-defined assertions, and the computational structure of the ML workflow. We show that the standard Thompson sampling approach is insufficient for switching models in ML workflows. In contrast, our proposed approaches are effective and scales to complex real-world ML workflows. Experiments show that switching models at runtime while reusing intermediate results provides higher accuracy, but also 48% efficiency gain compared to sequential workflow runs.