Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge
基于LLM作为评判者的5G领域知识与故障分析大模型自由文本评估
机构 * Surrey Institute for People-Centered Artificial Intelligence(萨里以人为本人工智能研究院) ; Google(谷歌)
专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL、cs.AI
AI总结 本文以自由文本格式评估Claude-Haiku-4.5等三个轻量型LLM的5G领域知识与故障分析能力,发现其故障诊断准确率超90%但规范召回不足,Gemini-3.1-Flash-Lite效率最优适合生产部署。
Comments 6pages, 4figures. Accepted for presentation in IEEE CSCN conference