arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22621cs.AI

Opti-Q:一种用于多语言模型问题规划的基于约束的优化框架

Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning

  • University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
  • Portland State University(波特兰州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Aamir Hamid, Bharg Barot, Satvik Racharla, Tim Finin, Primal Pappachan, Roberto Yus

AI总结:

研究针对多语言模型预算部署难题,提出OPTI-Q框架。该框架将模型调用建模为执行DAG操作符,利用PERFDB估计质量和成本,经帕累托前沿搜索选计划,在指定预算下于MMLU-Pro和SimpleQA上提升QoA,实现更好的质量-资源权衡。

AI中文摘要:

虽然大语言模型(LLMs)能够实现强大的问答(QA)功能,但预算部署因不确定性和异构资源配置文件(成本、延迟和能源)而变得复杂。我们提出了OPTI-Q,这是一个受数据库启发、基于成本的优化器,它为多语言模型编排实现了执行前规划范式。OPTI-Q将语言模型调用建模为执行DAG中的物理操作符,并针对每个问题搜索在用户指定的资源约束下权衡财务成本、延迟和能源的同时优化答案质量(QoA)的计划。计划可以包括将中间答案作为上下文传递的顺序操作符以及并发运行模型并合并其输出的并行/混合操作符。为了在不执行每个候选计划的情况下搜索这个空间,OPTI-Q使用PERFDB,一个从基准测试和执行跟踪中填充和刷新的统计目录,来估计单个操作符和组合子计划的QoA和资源成本。使用这些估计值,OPTI-Q执行帕累托前沿搜索并根据用户偏好选择最终计划。在用户指定预算下的MMLU-Pro和SimpleQA上,OPTI-Q在可比成本下比基线平均提高了约58%和约41%的QoA,表明数据库式规划为多语言模型QA带来了更好的质量-资源权衡。

英文摘要:

While large language models (LLMs) enable strong question answering (QA), budgeted deployment is complicated by nondeterminism and heterogeneous resource profiles (cost, latency, and energy). We present OPTI-Q, a database-inspired, cost-based optimizer that implements a plan-before-execute paradigm for multi-LLM orchestration. OPTI-Q models LLM invocations as physical operators in an execution DAG and, for each question, searches for plans that optimize answer quality (QoA) while trading off financial cost, latency, and energy under user-specified resource constraints. Plans can include sequential operators that pass intermediate answers as context and parallel/blend operators that run models concurrently and merge their outputs. To search this space without executing each candidate plan, OPTI-Q uses PERFDB, a statistics catalog populated and refreshed from benchmarks and execution traces, to estimate the QoA and resource costs of both individual operators and composed subplans. Using these estimates, OPTI-Q performs Pareto-frontier search and selects a final plan based on user preferences. On MMLU-Pro and SimpleQA under user-specified budgets, OPTI-Q improves average QoA by ~58% and ~41% over baselines at comparable cost, demonstrating that database-style planning yields better quality-resource trade-offs for multi-LLM QA.

↑