arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HarnessEval-W:将视觉世界评估智能体化

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

Weiliang Chen, Haowen Sun, Jun Gao, Jiawei Chi, Hanyang Wang, Qiyu Dai, Yihao Li, Hao Li, Jingnan Gao, Yi-Hsin Hung, Xingzhuo Guo, Shangchen Miao, Zhiyuan Shi, Xiang Li, Fengrui Tian, Weihua Du, Ziqi Huang, Shenyuan Gao, Siqiao Huang, Mingyu Liu, Yifei Li, Shizun Wang, Xi Wang, Tianqi Zhang, Xue Luo, Xiyin Ren, Jinshan Ren, Xiaoyang Shen, Xiaobo Hu, Zhiyang Dou, Mingyu Ding, Yichao Yan, Xinchao Wang, Yizhou Wang, Shilong Liu, Wenzhao Zheng, Yueqi Duan, Yuan Gong, Ziwei Liu, Ming-Yu Liu, Jialong Wu, Jiangran Lyu, Fangfu Liu

arXiv 2608.16859首次发表:更新:

AI 中文总结

该研究提出智能体化评估流程HarnessEval-W,将LLM生态的harness范式应用于世界模型基准测试,对18个世界模型的330个案例评估,其判断与人类偏好一致且提供可验证诊断,已开源并邀社区贡献。

AI 中文摘要

一个基准测试不应仅提供标量分数,使评估值得信赖的是证明该分数合理的推理过程。对于世界模型而言,这一点尤为关键,因为判断其生成的序列(rollout)需要理解物理、因果关系和世界状态是否正确演化。人类能自然发现此类违规情况,但现有基准测试均未实现这一能力的自动化:指标通过蛮力计算得出,未留下可检查或验证的推理链。我们推出HarnessEval-W,这是一种智能体化评估流程,将LLM生态系统中的harness范式应用于世界模型基准测试。HarnessEval-W不采用固定规则,而是解读每个评估案例的上下文,将评估问题分解为可测量的子问题,并生成专门的子智能体,每个子智能体配备定制化上下文和诊断工具,用于推理自身的子问题。随后,父智能体会验证收集到的证据,并将其总结为最终判断。这种分层工作流程将每次评估转化为透明的证据树,其完整推理链为结果提供支撑。我们将HarnessEval-W应用于18个代表性世界模型的330个评估案例。其判断与人类偏好高度一致,同时为每个生成的序列提供可验证的细粒度诊断。我们将完整流程作为实时基准测试开源,并邀请广大社区随着世界模型的发展贡献新技能和评估案例。

英文摘要

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.

CommentsProject Page: https://mirros-lab.github.io/HarnessEval-W

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑