Do You Get the Hint? Benchmarking LLMs on the Board Game Concept
你得到提示了吗?对大型语言模型的棋盘游戏概念基准测试
机构 * CLiPS, University of Antwerp(CLiPS,安特卫普大学)
专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL
AI总结 本文通过棋盘游戏Concept测试LLMs的演绎推理能力,发现人类易解而LLMs表现不佳,尤其在低资源语言中更差。