(Towards) Scalable Reliable Automated Evaluation with Large Language Models
面向可扩展的可靠自动化评估:基于大语言模型
机构 * KIT(卡尔斯鲁厄理工学院)
AI总结 本研究提出一种基于多LLM两两比较与Elo评分的自动化评估框架,可近似专家评估,减少人工干预,实现对LLM输出的高效可靠评估。
Comments 17 pages. Published in the Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM2) at ACL 2025
Journal ref Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM2), pages 320-336, Association for Computational Linguistics, 2025