面向云计算环境中LLM驱动文本分类的基于工作池编排的智能体自动扩缩容
Agentic Autoscaling through Worker-Pool Orchestration for LLM-driven Text Classification in Cloud Computing Environments
浏览论文内容
中文总结 AI 辅助
针对LLM文本分类的高延迟与突发负载,提出基于工作池编排的智能体自动扩缩容框架,集成优先级队列与动态工作池,在AG News和SMS Spam数据集上分别达到90.5%和99.5%准确率,优于基线并提升资源效率。
中文摘要 AI 辅助
基于大语言模型(LLM)的系统在大规模文本处理中的日益普及,产生了对动态自动扩缩容的迫切需求,以管理高延迟、突发性和计算密集型工作负载。本文提出了一种通过工作池编排实现LLM驱动文本分类的智能体自动扩缩容框架。该框架集成了优先级任务队列、动态智能体工作池、实时指标收集器和应用层自动扩缩容器。其分类器无关的设计支持零样本和微调语言模型,而无需修改自动扩缩容逻辑。该框架使用Autoscaling+BART和Autoscaling+DeBERTa进行评估,并与静态分配以及独立的RoBERTa和DistilBERT基线进行比较。在AG News数据集上,Autoscaling+BART达到84.5%的准确率,而Autoscaling+DeBERTa将其提升至90.5%。在SMS Spam Collection数据集上,Autoscaling+DeBERTa达到99.5%的准确率,而Autoscaling+BART以较低的执行时间达到84.5%的准确率。总体而言,所提出的框架在资源效率方面始终优于基线方法,同时保持高分类性能,表明弹性工作池编排为云环境中可扩展的LLM驱动文本分类提供了一种有效且成本高效的解决方案。
英文摘要
The growing adoption of large language model (LLM)-based systems for large-scale text processing has created a critical need for dynamic autoscaling to manage high-latency, bursty, and computationally intensive workloads. This paper proposes an agentic autoscaling framework through worker-pool orchestration for LLM-driven text classification. The framework integrates a priority task queue, a dynamic pool of agent workers, a real-time metrics collector, and an application-layer autoscaler. Its classifier-agnostic design supports both zero-shot and fine-tuned language models without modifying the autoscaling logic. The framework is evaluated using Autoscaling+BART and Autoscaling+DeBERTa against static allocation and standalone RoBERTa and DistilBERT baselines. On the AG News dataset, Autoscaling+BART achieves 84.5% accuracy, while Autoscaling+DeBERTa improves it to 90.5%. On the SMS Spam Collection dataset, Autoscaling+DeBERTa achieves 99.5% accuracy, whereas Autoscaling+BART attains 84.5% accuracy with lower execution time. Overall, the proposed framework consistently outperforms the baseline approaches in resource efficiency while maintaining high classification performance, demonstrating that elastic worker-pool orchestration provides an effective and cost-efficient solution for scalable LLM-driven text classification in cloud environments.
发表机构
- The University of Melbourne(墨尔本大学)
- Banaras Hindu University(贝拿勒斯印度教大学)
机构由 AI 辅助整理,请以论文原文为准。