arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ScholarStack:面向科学智能体的分层研究资产编排与跨任务复用

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

ScholarSeed AI Team, Caoqinwei Gong, Xue Jiang, Wei Luo, Xiaoyu Qiu, Jiayi Sheng, Yi Wang, Zheng Yu, Ao Zhang, Haifan Zhang, Hanwei Zhang, Jihai Zhang, Yuan Cao, Wei Chen, Liyun Dai, Wenkai Fang, Guanglei Wang, Kai Ying, Tingyu Zhu, Wotao Yin

arXiv 2609.23735首次发表:更新:

发表机构

DAMO Academy, Alibaba Group(阿里巴巴集团达摩院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ScholarStack分层研究资产框架,将论文编译为可复用资产,通过跨任务复用提升科学智能体在需跨论文证据任务上的质量并降低查询成本。

AI 中文摘要

科学智能体支持一系列基于文献的研究任务,如检索、问答、基于证据的生成和主张评估。然而,现有大多数系统都是围绕单个任务组织的:相同的论文被反复检索、分段和解读,在一个任务中构建的理解难以在下一个任务中复用。我们提出ScholarStack,一个分层研究资产框架,它将论文集合编译成可复用、带版本且保留来源的资产,涵盖三个互补层级:基于来源的论文级陈述、领域级组织以及基于证据的跨论文综合。一个通用访问接口以每个任务所需的证据粒度返回任务特定视图,保留研究条件、来源可追溯性和验证状态。我们在跨越十个任务设置的四个任务族上实例化该框架,在匹配的基础模型下,将使用编译资产的智能体与任务特定基线进行比较。质量提升集中在需要跨论文证据的任务上,如多论文问答和文献综述生成,并且在所有测量了查询时令牌成本的任务上,该成本均有所下降,因为资产被编译一次并在多个任务中复用。这些结果表明,分层研究资产可以作为科学智能体的共享基础设施,将基于文献的辅助从孤立的文档处理转向累积的、基于证据的工作流。

英文摘要

Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however, are organized around individual tasks: the same papers are repeatedly retrieved, segmented, and interpreted, and the understanding built in one task is difficult to reuse in the next. We present ScholarStack, a layered research asset framework that compiles a paper collection into reusable, versioned, and provenance-preserving assets at three complementary levels: source-grounded paper-level statements, domain-level organization, and evidence-grounded cross-paper syntheses. A common access interface returns task-specific views at the evidence granularity each task requires, preserving study conditions, source traceability, and verification status. We instantiate the framework on four task families spanning ten task settings, comparing agents that use the compiled assets with task-specific baselines under matched base models. Quality gains concentrate on tasks that require cross-paper evidence, such as multi-paper question answering and literature review generation, and query-time token cost falls on every task where it is measured, with assets compiled once and reused across tasks. These results suggest that layered research assets can serve as shared infrastructure for scientific agents, shifting literature-based assistance from isolated document processing toward cumulative, evidence-grounded workflows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑