arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于光网络自动化的大语言模型的人工基础评估

Human Grounded Evaluation of Large Language Models for Optical Network Automation

Kiarash Rezaei, Omran Ayoub, Paolo Monti, Carlos Natalino

arXiv 2607.18068首次发表:更新:

发表机构

University of Applied Sciences and Arts of Southern Switzerland(瑞士南瑞士应用科学与艺术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对光网络自动化中LLMs输出质量和成本差异大的问题,提出HuGLEN评估管道,结合LLM评判器和专家评级,用QES排名,用于转换XAI模型输出,结果显示中等规模LLM的QES最高,减轻人工标注负担,助力模型选择。

AI 中文摘要

大语言模型(LLMs)越来越多地用于网络自动化,但其输出质量和推理成本在不同LLM系列中差异很大。我们提出了HuGLEN,这是一个逐步评估管道,使用LLM作为评判器并结合少量专家评级,以实现对候选LLMs的可扩展和可重复比较,并使用质量效率得分(QES)对它们进行排名。我们展示了HuGLEN将可解释人工智能(XAI)模型用于光网络传输质量(QoT)估计任务的输出转换为操作员友好的解释。结果表明,一个中等规模的LLM(12B参数)实现了最高的QES,表明在解释质量和效率之间取得了最佳平衡。总体而言,HuGLEN减轻了人工标注负担,同时支持面向操作员的自动化任务的一致模型选择。

英文摘要

Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an LLM-as-a-judge together with a small set of expert ratings to enable scalable and reproducible comparison of candidate LLMs, and to rank them using a quality efficiency score (QES). We demonstrate HuGLEN for translating outputs from an explainable artificial intelligence (XAI) model for the optical network quality of transmission (QoT) estimation task into operator-friendly explanations. Our results show that a medium-sized LLM (12B parameters) achieves the highest QES, indicating the best trade-off between explanation quality and efficiency. Overall, HuGLEN reduces the human-labeling burden while supporting consistent model selection for operator-facing automation tasks.

DOI:10.1109/ICTON71309.2026.11704905}

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑