EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
EnvSimBench:一个评估和改进基于LLM的环境模拟的基准
Yi Liu, TingFeng Hui, Wei Zhang, Li Sun, Ningxin Su, Jian Wang, Sen Su
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Chongqing University(重庆大学)
CommentsWe have further refined the benchmark construction and reference verification pipeline to improve clarity and consistency. The revised version includes updated results and additional details to better align the evaluation with the intended setup. These changes provide a more precise presentation of the experimental findings, with conclusions and contributions remaining unchanged
机构
*
RMIT University(皇家墨尔本理工大学)
;
Monash University(墨尔本大学)
;
Adelaide University(阿德莱德大学)
;
The University of Hong Kong(香港大学)
;
ESPOL University(ESPOL大学)
TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering
机构
*
Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳))
;
City University of Macau, Macao SAR, China(澳门城市大学)
;
Peng Cheng Laboratory, Shenzhen, China(鹏城实验室)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL
RAG or Learning? Understanding the Limits of LLM Adaptation under Continuous Knowledge Drift in the Real World
RAG还是学习?在现实世界中连续知识漂移下理解LLM适应的极限
Hanbing Liu, Lang Cao, Yang Li
机构
*
Tsinghua University(清华大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
School of Artificial Intelligence, Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)人工智能学院)
;
Shenzhen Key Laboratory of Ubiquitous Data Enabling, Tsinghua Shenzhen International Graduate School, Tsinghua University(深圳 ubiquitous 数据赋能重点实验室,清华大学深圳国际研究生院,清华大学)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);pretraining(abstract)
The Catastrophic Paradox of Human Cognitive Frameworks in Large Language Model Evaluation: A Comprehensive Empirical Analysis of the CHC-LLM Incompatibility
人类认知框架在大语言模型评估中的灾难性悖论:对CHC-LLM不兼容性的全面实证分析
Mohan Reddy
机构
*
Stanford University(斯坦福大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(title);分类 cs.AI
LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical Study
Dongil Yang, Minjin Kim, Sunghwan Kim, Beong-woo Kwak, Minjun Park, Jinseok Hong, Woontack Woo, Jinyoung Yeo
机构
*
Department of Artificial Intelligence, Yonsei University(人工智能系,延世大学)
;
Graduate School of Metaverse, KAIST(元宇宙研究生院,韩国科学技术院)
;
Graduate School of Culture Technology, KAIST(文化科技研究生院,韩国科学技术院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(title);分类 cs.CL
机构
*
The International Joint Institute of Tianjin University(天津大学国际联合研究院)
;
Tianjin University(天津大学)
;
TJUNLP Lab, School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院)
;
National Governance Institute, Tianjin Normal University(天津师范大学国家治理研究院)
;
College of Computer and Information Engineering, Tianjin Normal University(天津师范大学计算机与信息工程学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI