发表机构
Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型推理框架代码复杂、维护成本高的问题,提出MetaInfer方法,通过用户指定运行时约束,利用大语言模型驱动多智能体协作系统结合契约知识库自动生成定制推理框架,并从多视角评估,实现从显式知识生成可运行方案。
AI 中文摘要
随着大语言模型技术的发展,模型家族、计算硬件、量化方案、并行化策略和专门的优化内核的范围不断扩大,大幅增加了通用推理框架的代码复杂性和维护成本。传统软件工程使用多层抽象来支持不同的应用场景,但这些抽象也增加了系统复杂性并可能引入额外的性能开销。本文提出了MetaInfer,一种“大语言模型即编译器”的方法,用户只需指定推理程序的运行时约束。一个由大语言模型驱动的多智能体协作系统,结合契约知识库,然后自动生成一个满足这些约束的紧凑定制推理框架。我们从三个角度评估MetaInfer:源代码引用的效果、在零引用约束下为CKB覆盖目标生成的引擎的运行时行为和性能概况,以及针对新模型和平台场景的知识库演变。结果表明,MetaInfer将生成约束、验证反馈和知识整合组织成一个连续的闭环,能够从显式知识中生成可运行的定制推理解决方案。代码可在这个https网址公开获取。
英文摘要
Large language models now evolve faster than production inference systems can be ported and optimized. New releases change attention, MoE routing, quantization formats, KV-cache layout, and parallel execution patterns, while deployed accelerator fleets remain heterogeneous across hardware generations, framework forks, operator libraries, compiler backends, and communication runtimes. Serving a new model on existing hardware is therefore a model-framework-kernel-hardware co-adaptation problem. We present MetaInfer, an LLM-agent system that formulates inference adaptation as route search over a costed execution-adaptation graph. The graph connects model semantics, framework dispatch, kernel choices, hardware capabilities, runtime evidence, and serving objectives. MetaInfer constructs and updates this graph during execution, restores missing or blocked routes through patches, and reduces route cost through staged validation and end-to-end profiling. Three real episodes -- DeepSeek V4 Flash on NVIDIA A800, GLM 5.3 Flash on NVIDIA A800, and DeepSeek V4 Flash on Hygon K100AI DCU -- demonstrate deployment repair, cross-model knowledge transfer, and portability across heterogeneous accelerator software stacks.