arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00785cs.MA

HIERA:面向内容发现系统的分层多智能体相关性评估框架

HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems

Pritom Saha Akash, Phanideep Gampa, Chao Shen, Ying Chen, Sheikh Muhammad Sarwar

AI总结:

该研究针对内容发现系统人工标注的问题,提出分层多智能体相关性评估框架HIERA,通过四个专用智能体的分层协调,在五个数据集上较11个基线方法取得显著性能提升。

AI中文摘要:

内容发现系统依赖相关性判断来评估搜索质量,但人工标注存在标注者间分歧和扩展性成本问题。虽然大型语言模型作为自动评估器展现出潜力,但现有方法依赖平面聚合策略:单步提示、投票集成或未协调的多智能体流水线,仅聚合独立输出而无整合。我们提出HIERA,这是一个分层多智能体相关性评估框架,包含四个专用智能体:相关性判断智能体、查询分析智能体、项目分析智能体和关系分析智能体。判断智能体确定何时需要专业分析;关系分析智能体随后协调查询与项目分析,并结合外部知识建立相关性关系以完成最终判断。消融研究显示,使用相同智能体和外部知识但无分层协调时性能下降,证实协调结构本身是性能提升的原因。在五个数据集(EVS、MSRD、ESCI、WANDS、Home Depot)上的评估显示,其优于11个基线方法:Home Depot数据集提升10.2%,ESCI数据集提升4.8%,EVS数据集提升高达38%(p<0.05);分层协调相较于使用相同智能体的未协调协作,提升12.7%。

英文摘要:

Content discovery systems depend on relevance judgment for search quality evaluation, but human annotation faces inter-annotator disagreement and scaling costs. While Large Language Models show promise as automated assessors, current approaches rely on flat aggregation strategies: single-step prompting, voting ensembles, or uncoordinated multi-agent pipelines that aggregate independent outputs without integration. We propose HIERA, a hierarchical multi-agent relevance assessment framework with four specialized agents: a Relevance Judge, Query Analyzer, Item Analyzer, and Relation Analyzer. The Judge determines when specialist analysis is needed; the Relation Analyzer then coordinates query and item analyses with external knowledge to establish relevance relationships for final judgment. Ablation studies show that the same agents and external knowledge without hierarchical coordination degrade performance, confirming that the coordination structure itself accounts for the improvement. Evaluation across five datasets (EVS, MSRD, ESCI, WANDS, Home Depot) shows improvements over 11 baselines: 10.2\% on Home Depot, 4.8\% on ESCI, and up to 38\% on EVS ($p < 0.05$). Hierarchical coordination yields 12.7\% improvement over uncoordinated collaboration using identical agents.

↑