发表机构
University of Applied Sciences and Arts of Southern Switzerland(瑞士南瑞士应用科学与艺术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对光网络自动化中LLMs输出质量和成本差异大的问题,提出HuGLEN评估管道,结合LLM评判器和专家评级,用QES排名,用于转换XAI模型输出,结果显示中等规模LLM的QES最高,减轻人工标注负担,助力模型选择。
AI 中文摘要
大语言模型(LLMs)越来越多地用于网络自动化,但其输出质量和推理成本在不同LLM系列中差异很大。我们提出了HuGLEN,这是一个逐步评估管道,使用LLM作为评判器并结合少量专家评级,以实现对候选LLMs的可扩展和可重复比较,并使用质量效率得分(QES)对它们进行排名。我们展示了HuGLEN将可解释人工智能(XAI)模型用于光网络传输质量(QoT)估计任务的输出转换为操作员友好的解释。结果表明,一个中等规模的LLM(12B参数)实现了最高的QES,表明在解释质量和效率之间取得了最佳平衡。总体而言,HuGLEN减轻了人工标注负担,同时支持面向操作员的自动化任务的一致模型选择。
英文摘要
Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an LLM-as-a-judge together with a small set of expert ratings to enable scalable and reproducible comparison of candidate LLMs, and to rank them using a quality efficiency score (QES). We demonstrate HuGLEN for translating outputs from an explainable artificial intelligence (XAI) model for the optical network quality of transmission (QoT) estimation task into operator-friendly explanations. Our results show that a medium-sized LLM (12B parameters) achieves the highest QES, indicating the best trade-off between explanation quality and efficiency. Overall, HuGLEN reduces the human-labeling burden while supporting consistent model selection for operator-facing automation tasks.
DOI:10.1109/ICTON71309.2026.11704905}