发表机构
University of Tsukuba; The University of Tokyo; Hiroshima University(筑波大学; 东京大学; 广岛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
JuryFlow提出分歧引导的人机协同多智能体评估框架,将评判分歧作为信号,通过图传播和规则归纳提升评估准确性。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被部署为AI生成内容的自动评判者,然而单一评判者不可靠,即使是一组评判者也会留下一个难题:当评判者意见不一致时,多数投票会丢弃冲突而不是解决它。我们提出了JuryFlow,一个分歧引导的人机协同多智能体评估框架,它将评判者间的分歧视为精确的、声明级别的信号,指示评估中哪里存在不确定性,而不是作为可平均化的噪声。JuryFlow将每个候选响应分解为原子声明,由一组异构评判者对每个声明进行裁决,并构建一个分歧图,其节点通过裁决熵进行评分,边编码声明之间的结构相似性。人类充当结构引导者,通过单一、最小干预选择要解决的分歧,而不是重新标记响应,之后焦点声明被重新评估,修正沿图边传播到历史相似案例,并被固化为所有评判者继承的可复用规则条目,使评估器逐步自我完善。为了在没有人类研究的情况下实现大规模、可复现的基准测试,我们在自动配置中评估JuryFlow,其中焦点选择由熵排名完成。在MT-Bench和LLMBar上,与单一评判者和多数投票小组基线相比,JuryFlow提高了与黄金标签的一致性,消融研究分离了分歧定向重评估、传播和规则归纳的贡献。我们贡献了(1)一种人机协同范式,将人类从标记者重塑为结构引导者,(2)JuryFlow框架,通过分歧图、焦点重评估和闭环规则归纳实现该范式,以及(3)一个带有消融的评估协议,以隔离增益的来源。
英文摘要
Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet a single judge is unreliable and even a panel of judges leaves a hard residue: when judges disagree, majority voting discards the conflict instead of resolving it. We present JuryFlow, a disagreement-guided, human-in-the-loop multi-agent evaluation framework that treats inter-judge disagreement not as noise to be averaged away, but as a precise, claim-level signal indicating where an evaluation is uncertain. JuryFlow decomposes each candidate response into atomic claims, has a panel of heterogeneous judges assign per-claim verdicts, and builds a disagreement graph whose nodes are scored by verdict entropy and whose edges encode structural similarity between claims. A human acts as a structural guide, selecting which disagreement to resolve through a single, minimal intervention rather than re-labeling the response, after which the focal claim is re-evaluated, the correction propagates along graph edges and to historically similar cases, and is crystallized into reusable rubric entries that all judges inherit, making the evaluator progressively self-refining. To enable large-scale, reproducible benchmarking without human studies, we evaluate JuryFlow in an automatic configuration in which focal selection is made by entropy ranking. On MT-Bench and LLMBar, JuryFlow improves agreement with gold labels over single-judge and majority-vote panel baselines, and ablations isolate the contributions of disagreement-targeted re-evaluation, propagation, and rubric induction. We contribute (1) a human-in-the-loop paradigm that recasts the human from labeler to structural guide, (2) the JuryFlow framework operationalizing it through a disagreement graph, focal re-evaluation, and closed-loop rubric induction, and (3) an evaluation protocol with ablations that isolate where the gains originate.
Comments9 pages, 5 figures, 5 tables. To appear in Proceedings of the 14th International Conference on Human-Agent Interaction (HAI '26), November 16-19, 2026, Osaka, Japan