AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
AgentHallu: 评估基于大语言模型的代理的自动幻觉归因
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; Department of Computer Science & Technology, Tsinghua University(清华大学计算机科学与技术系) ; University of California, Santa Barbara(加州大学圣芭芭拉分校) ; Center for Research on Intelligent Perception and Computing, NLPR, CASIA(智能感知与计算研究中心,国家工程实验室)
专题命中 规划推理 :reasoning(abstract);planning(abstract);分类 cs.CL
AI总结 AgentHallu提出一个评估基于大语言模型的代理自动识别幻觉来源的任务,通过高质量轨迹和多级注释评估13个模型,发现顶级模型在定位幻觉步骤上表现有限,工具使用幻觉最难识别。
Comments Project page: https://liuxuannan.github.io/AgentHallu.github.io/