发表机构
Southern Illinois University Carbondale(南伊利诺伊大学卡本代尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对边缘-云LLM系统中函数调用的碳排放问题,提出碳感知路由框架,利用k-NN预测器结合实时碳强度,将查询路由至最低排放层级,在保持云端准确性的同时平均减少4倍碳排放。
AI 中文摘要
具有函数调用能力的大型语言模型(LLMs)正成为现代智能体AI系统的关键。然而,当前的部署通常将推理路由到强大的云端模型,导致显著的能源使用和碳排放。我们通过一个碳感知路由框架来应对这一可持续性挑战,该框架在三级边缘-云架构中分发函数调用查询,结合了异构硬件上的边缘和云端LLMs。其核心是一个轻量级k-NN预测器,在统一的语义-词汇嵌入空间中操作,估计每个边缘层级上查询特定的准确性、延迟和功耗。这些预测随后与实时电网碳强度相结合,将每个查询路由到能够成功执行它的最低排放层级。在最先进的函数调用基准和LLM系列上评估,我们的框架在匹配云端准确性的同时,平均将运营碳排放减少了4倍。
英文摘要
Large Language Models (LLMs) with function-calling capabilities are becoming critical for modern agentic AI systems. Nevertheless, current deployments typically route inferences to powerful cloud-based models, incurring significant energy use and carbon emissions. We address this sustainability challenge with a carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture, combining edge and cloud LLMs on heterogeneous hardware. At its core, a lightweight k-NN predictor operating in a unified semantic-lexical embedding space estimates query-specific accuracy, delay, and power consumption on each edge tier. These predictions are then combined with real-time grid carbon intensity to route every query to the lowest-emission tier capable of executing it successfully. Evaluated on state-of-the-art function-calling benchmarks and LLM families, our framework matches cloud-level accuracy while reducing operational carbon emissions by $4\times$ on average.
Comments2026 IEEE 33rd International Conference on Electronics, Circuits and Systems (ICECS)