arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JOVE:面向资源感知的LLM任务图的联合执行与验证

JOVE: Joint Execution and Verification for Resource-Aware LLM Task Graphs

Haoran Zhang, Dongjun Kim, Seohyeon Cha, Kevin S Chan, Ananthram Swami, Gustavo De Veciana, Haris Vikalo

arXiv 2610.03296首次发表:更新:

发表机构

University of Texas at Austin; DEVCOM Army Research Laboratory(德克萨斯大学奥斯汀分校; DEVCOM陆军研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

JOVE是一个在线框架,通过联合分配执行LLM和选择验证输出,在预算和延迟约束下优化推理任务图的执行与学习权衡,显著降低成本与延迟。

AI 中文摘要

复杂的推理查询可以被分解为有向无环任务图,并分布在异构的大语言模型(LLM)之间,通过并行性减少延迟,并使较小的模型能够解决复杂的任务。然而,在实践中,一个LLM对于给定子任务的适用性可能是先验未知的,仅执行并不能揭示输出的正确性。我们提出了JOVE,一个在线框架,它联合分配执行LLM并选择中间输出进行付费验证。验证异步运行,并用于改进未来的分配,因此系统必须在当前执行支出与后续学习之间取得平衡。我们研究了如何在长期预算和每次查询延迟约束下优化这种权衡,其中LLM服务质量、调用成本和执行时间是随机的且初始未知。JOVE通过求解一系列每次查询的混合整数线性规划来做出执行和验证决策。在线学习根据验证反馈更新任务相关的LLM质量估计,而信息增益奖励将学习的价值纳入分配决策。在一组自然假设下,我们为JOVE建立了次线性的质量学习遗憾。在四个推理基准上,JOVE在达到与标准推理基线相当的准确性的同时,将平均成本和延迟降低了至少3.17倍。

英文摘要

Complex reasoning queries can be decomposed into directed acyclic task graphs and distributed across heterogeneous LLMs, reducing latency through parallelism and enabling smaller models to solve complex tasks. In practice, however, the suitability of an LLM for a given subtask may be a priori unknown, and execution alone does not reveal output correctness. We propose JOVE, an online framework that jointly assigns executor LLMs and selects intermediate outputs for paid verification. Verification runs asynchronously and is used to improve future allocations, so the system must balance spending on execution now against learning for later. We study how to optimize this trade-off under a long-term budget and a per-query latency constraint, with stochastic, initially unknown LLM service quality, invocation costs, and execution times. JOVE makes execution and verification decisions by solving a sequence of per-query mixed-integer linear programs. Online learning updates task-dependent estimates of LLM quality based on verification feedback, while an information-gain bonus incorporates the value of learning into allocation decisions. Under a natural set of assumptions, we establish sublinear quality-learning regret for JOVE. Across four reasoning benchmarks, JOVE achieves competitive accuracy against standard inference baselines while reducing average cost and latency by at least 3.17 times.

Commentspreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑