Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
Agent-ValueBench: 一个全面的评估代理价值观的基准
机构 * State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室) ; School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) ; Peking University School of Software and Microelectronics(北京大学软件与微电子学院) ; Peking University School of Psychological and Cognitive Sciences(北京大学心理与认知科学学院) ; Key Laboratory of Machine Perception (Ministry of Education), Peking University(北京大学机器感知重点实验室)
专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.AI
AI总结 本文提出Agent-ValueBench,首个专门评估代理价值观的基准,包含394个环境和4335个价值冲突任务,揭示代理价值观与基础LLM的差异及Harness和技能对价值观的影响。