A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
语言模型智能体中目标导向性的行为与表征评估
机构 * University of Pennsylvania(宾夕法尼亚大学) ; New York University(纽约大学) ; Indiana University, Bloomington(印第安纳大学,布卢明顿) ; Northeastern University(东北大学) ; University College London(伦敦大学学院)
专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出一种结合行为评估与内部表征可解释性分析的目标导向性评估框架,并以LLM智能体在2D网格世界中的导航为例,验证了其行为与表征的一致性。
Comments Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)