SEAGym: An Evaluation Environment for Self-Evolving LLM Agents
SEAGym: 自我进化LLM智能体的评估环境
Congjie Zheng, Chuanyi Xue, Bin Liang, Jun Yang, Changshui Zhang
机构
*
Department of Automation, Tsinghua University(清华大学自动化系)
;
Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University(北京信息科学与技术国家研究中心(BNRist),清华大学)
机构
*
School of Computer Science, Chongqing University(重庆大学计算机学院)
;
AI Research Institution, Mashang Financial Institution(马上金融人工智能研究院)
;
Department of Information, Third Military Medical University(陆军军医大学信息系)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
Comments7 pages, 1 figure. Accepted for publication in the Main Research Track of the Twenty-First International Conference on Software Engineering Advances (ICSEA 2026)
The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025
CommentsPreregistered study (OSF: 10.17605/OSF.IO/U5NEY); deviations from the registered plan, including a formal validation-gate breach, are disclosed in Section 4.6. Companion study: arXiv:2606.29540. 3 figures, 4 tables
Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
多语言环境和低资源语言中LLM-as-a-Judge的挑战与建议
A. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
机构
*
LT3, IDLab, Universiteit Gent, Barcelona Supercomputing Center, LMU Munich & Munich Center for Machine Learning, German Center for Addiction Research in Childhood and Adolescence, University Medical Center Hamburg-Eppendorf, Mila - Quebec AI Institute, McGill University, Canada CIFAR AI Chair(LT3、IDLab、根特大学、巴塞罗那超级计算中心、慕尼黑莱茵河大学及慕尼黑机器学习中心、德国成年期成瘾研究中心、汉堡埃彭多夫大学医学中心、魁北克人工智能研究所、麦吉尔大学、加拿大 CIFAR 人工智能主席)