arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20777cs.CL

关注点树:面向科学评论中未陈述局限性提取的分层多智能体辩论

Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique

Sahil Mishra, Niranjan Rajeev, Tanmoy Chakraborty

首次发表
浏览论文内容

中文总结 AI 辅助

针对科学论文未陈述局限性提取的需求,本文提出Tree-of-Concerns多智能体框架,经ToC-Bench实验验证,其较最强基线提升79%准确率与11%覆盖度,可辅助审稿人开展系统性评估。

中文摘要 AI 辅助

随着科学文献数量增长,且论文愈发倾向于少报告自身局限性,多智能体大语言模型(LLM)为系统揭示这些隐藏的失效模式提供了有前景的方法。本文提出Tree-of-Concerns(关注点树),这一多智能体框架部署了具备类别特定分析视角的专业化质疑角色,作为并行辩论树,用于从科学论文中提取未陈述的局限性。每个角色开展结构化、基于证据的论证,而Panel Review(评审团)机制从全部五个视角重新评估每个保留的主张,以纠正类别漂移和严重程度校准偏差。通过在ToC-Bench(本文的基准数据集,包含414篇研究论文及1905个未陈述局限性,这些局限性源自审稿人报告的缺陷及后续引用评论)上的实验,我们证实,与最强基线相比,ToC将准确率提升了79%,覆盖度提升了11%,并能呈现具体、基于证据的关注点,以支持审稿人开展系统性评估。

英文摘要

As scientific literature grows and papers increasingly under-report limitations, multi-agent LLMs offer a promising approach to systematically uncover these hidden failure modes. Here, we introduce Tree-of-Concerns, a multi-agent framework that deploys specialized skeptic personas, each operating through a category-specific analytical lens, as parallel debate trees to extract unstated limitations from scientific papers. Each persona conducts structured, evidence-grounded argumentation, while a Panel Review mechanism re-evaluates each surviving claim from all five perspectives to correct category drift and severity miscalibration. Through retrieval-free, single-paper experiments on ToC-Bench, our benchmark of 414 research papers with 1,905 unstated limitations, sourced from reviewer-reported weaknesses and follow-up citation critiques, we demonstrate that ToC improves precision by 79% and coverage by 11% relative to the strongest baseline, surfacing specific, evidence-grounded concerns that support reviewers in systematic evaluation.

发表机构

  • IIT Delhi(印度理工学院德里分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑