arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21021cs.CLcs.AIcs.NI

基于LLM作为评判者的5G领域知识与故障分析大模型自由文本评估

Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

  • Surrey Institute for People-Centered Artificial Intelligence(萨里以人为本人工智能研究院)
  • Google(谷歌)

机构由 AI 辅助整理,请以论文原文为准。

Rishiraj Sengupta, Sotiris Chatzimiltis, Mohammad Shojafar, Xiatian Zhu

AI总结:

本文以自由文本格式评估Claude-Haiku-4.5等三个轻量型LLM的5G领域知识与故障分析能力,发现其故障诊断准确率超90%但规范召回不足,Gemini-3.1-Flash-Lite效率最优适合生产部署。

AI中文摘要:

5G及新兴6G网络中的实际故障分析需要领域专业知识来分析自由文本诊断内容,包括根本原因解释和推荐操作。大语言模型(LLM)已成为自动化此类工作的有前景方法,但轻量型、可边缘部署的模型是否能执行深入的自由文本诊断仍是未解决的问题。现有基准依赖带有固定答案的限制性多项选择题,而本文以自由文本生成格式评估5G领域理解与故障分析。转向该范式需要在开放式诊断推理上评估轻量型、可边缘部署的AI模型,同时需要可靠框架大规模验证这些文本输出。为此,本文在三个基准(TeleQNA ORAN FT、5G-Faults FT、TeleInter FT)的自由文本5G领域知识与故障分析任务上评估三个轻量型LLM:Claude-Haiku-4.5、GPT-5.4-Mini和Gemini-3.1-Flash-Lite。三名独立前沿评判者对输出打分,通过成对评判者间一致性衡量LLM作为评判者方法的有效性。所有三个模型在故障诊断上准确率至少达90%,但3GPP和O-RAN规范的零样本召回仍是关键差距,所有模型得分均低于60%。所有运行的平均评判者间一致性至少为0.90,表明多评判者LLM打分可为开放式电信响应产生一致、可复现的评分。从业务角度看,Gemini-3.1-Flash-Lite实现最佳效率权衡,兼具有竞争力的准确率、最低推理成本与延迟,是电信生产部署的最合适候选。

英文摘要:

Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause explanations and recommended actions. LLMs have emerged as a promising approach to automating this, yet whether lightweight, edge-deployable models are capable of performing in-depth free-text diagnostics remains an open question. While existing benchmarks rely on restrictive MCQs with fixed answer keys, this paper evaluates 5G domain understanding and fault analysis in a free-text generation format. Transitioning to this paradigm requires evaluating lightweight, edge-deployable AI models on open-ended diagnostic reasoning, alongside a dependable framework to validate these text outputs at scale. To address this we evaluate three lightweight LLMs, Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite, on free-text 5G domain knowledge and fault-analysis tasks across three benchmarks, TeleQNA ORAN FT, 5G-Faults FT, and TeleInter FT. Three independent frontier judges score outputs, and pairwise inter-judge agreement is measured as an empirical test of the LLM-as-Judge methodology. All three models reach at least 90% accuracy on fault diagnosis, while zero-shot recall of 3GPP and O-RAN specifications remains the critical gap, with all models scoring below 60%. Mean inter-judge agreement is at least 0.90 across all runs, indicating that multi-judge LLM scoring produces consistent, reproducible grades for open-ended telecom responses. Operationally, Gemini-3.1-Flash-Lite offers the best efficiency trade-off, combining competitive accuracy with the lowest inference cost and latency, making it the most suitable candidate for production telecom deployments.

补充信息

↑