arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估生成式AI的推理计算:面向企业工作负载的框架

Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads

Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid

arXiv 2610.07094首次发表:更新:

发表机构

Citigroup Inc.; Ernst & Young LLP; NVIDIA Corporation(花旗集团; 安永会计师事务所; 英伟达公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体轨迹推理,提出四层评估框架,以有效吞吐量和每次成功回合成本为核心指标,并论证确定性、尾延迟及每步可靠性对长时域任务性能的关键影响。

AI 中文摘要

LLM部署正从单轮补全转向智能体轨迹,在此过程中,模型在行动前于测试时进行规划、调用工具、读取结果并进行推理。这颠覆了推理硬件的经济学:聊天服务通过大批量分摊权重读取,而智能体轨迹是顺序依赖的,以有效批大小1运行,并使每令牌解码延迟(TPOT)成为任务完成时间中的主导项。通过roofline分析和封闭形式的回合延迟模型,我们展示了为何该机制青睐将权重保留在片上SRAM(Cerebras WSE-3/3T,Groq/NVIDIA LPU)或编译器管理的分层内存(SambaNova SN40L/SN50)中的加速器,以及为何三个供应商生态系统在2026年汇聚于分离式预填充/解码服务。我们表明,每步可靠性随轨迹长度呈指数级复合——2%的每步失败率会抹去20步智能体2倍的解码优势——因此确定性和尾延迟是一阶性能变量。随后,我们提出了一个四层评估框架(硅片、服务系统、智能体回合、企业),并基于智能体SLO下的有效吞吐量和每次成功回合的成本构建了指标集,一个覆盖六个任务族的六轴基准测试协议,一个配对自助统计设计,一个用于供应商运行基准测试的证明协议,以及具有明确盈亏平衡条件的TCO、可用性和采用时机模型。所有性能数据均为公开且按证据类别标注;我们陈述了七个可证伪假设及检验它们的实验,并论证最可能的原始结果是,在长时域工作上,令牌吞吐量排名与每次成功任务成本排名存在分歧。

英文摘要

LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentially dependent, run at effective batch one, and make per-token decode latency (TPOT) the dominant term in task completion time. Using a roofline analysis and a closed-form episode-latency model, we show why this regime favours accelerators that keep weights in on-die SRAM (Cerebras WSE-3/3T, Groq/NVIDIA LPU) or compiler-managed tiered memory (SambaNova SN40L/SN50), and why three vendor ecosystems converged in 2026 on disaggregated prefill/decode serving. We show that per-step reliability compounds exponentially in trajectory length-a 2% per-step failure rate erases a 2x decode advantage for a 20-step agent-so determinism and tail latency are first-order performance variables. We then propose a four-layer evaluation framework (silicon, serving system, agent episode, enterprise) with a metric set built on goodput at an agentic SLO and cost per successful episode, a six-axis benchmark protocol over six task families, a paired-bootstrap statistical design, an attestation protocol for vendor-run benchmarks, and TCO, availability and adoption-timing models with explicit break-even conditions. All performance figures are public and labelled by evidence class; we state seven falsifiable hypotheses and the experiments that test them, and argue that the most likely original result is that token-throughput rankings diverge from cost-per-successful-task rankings on long-horizon work.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑