发表机构
University of Alabama in Huntsville(阿拉巴马大学亨茨维尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Paper Pilot,一种应用科学领域可追溯证据的人在环专家系统,通过8个批准门等机制解决AI辅助手稿生成的治理问题,实证验证其可减少引用伪造,将LLM辅助写作定位为受控人机决策支持流程。
AI 中文摘要
大型语言模型(LLM)智能体正越来越多地嵌入科学工作流程,用于文献分析、起草和评审。现有系统推进了自主发现和手稿生成,但未解决治理问题:当想法、方法、结果和主张在AI辅助工作流程中传播时,缺乏强制性人类批准或人工制品级别的可追溯性。本文提出Paper Pilot,一种面向应用科学的可追溯证据的科学手稿生成的人在环专家系统。它通过手稿所有者批准门、明确的不通过标准、主张分类、审计日志、咨询式LLM评审和证据锁定修订控制,将协作智能体推理工程(Collaborative Agent Reasoning Engineering,CARE)方法适配到手稿开发中。该框架在从想法到主张的流程中定义了8个批准门,并区分基于文献的主张与基于人工制品的主张,要求报告的数字和解释必须可追溯至已批准的证据;其系统提示已公开发布,可部署在ChatGPT、Gemini、Claude或机构LLM环境中。作为首次实证验证,我们使用受控的机械评分基准(2个商业LLM、真实arXiv论文、无LLM评审员)评估基于引用的层:在覆盖范围压力下,无门控的起草者最多伪造25%的引用且从未标记证据缺口,而相同模型在Paper Pilot的证据锁定规则下,未产生伪造引用并将植入的缺口作为显式占位符呈现。结果可追溯性、修订和对抗鲁棒性的初步结果也指向相同方向;完整评估留待未来工作。Paper Pilot将LLM辅助写作定位为受控的人机决策支持流程,而非完全自主的创作管道。
英文摘要
Large language model (LLM) agents are increasingly embedded in scientific workflows for literature analysis, drafting, and review. Existing systems advance autonomous discovery and manuscript generation, but do not resolve the governance problem that arises when ideas, methods, results, and claims propagate through AI-assisted workflows without mandatory human approval or artifact-level traceability. This paper proposes Paper Pilot, a human-in-the-loop expert system for evidence-traceable scientific manuscript generation in applied sciences. It adapts the Collaborative Agent Reasoning Engineering (CARE) methodology to manuscript development through manuscript-owner approval gates, explicit no-pass criteria, claim classification, audit logging, advisory LLM review, and evidence-locked revision control. The framework defines eight approval gates across the idea-to-claim pipeline and distinguishes literature-grounded from artifact-grounded claims, requiring reported numbers and interpretations to remain traceable to approved evidence; its system prompt is openly released for deployment in ChatGPT, Gemini, Claude, or institutional LLM environments. As a first empirical validation, we evaluate the citation-grounding layer with a controlled, mechanically scored benchmark (two commercial LLMs, real arXiv papers, no LLM judge): under coverage pressure ungated drafters fabricated up to 25% of their citations and never flagged an evidence gap, whereas the same models under Paper Pilot's evidence-locked rules produced zero fabricated citations and surfaced the planted gaps as explicit placeholders. Preliminary results for result grounding, revision, and adversarial robustness point the same way; full evaluation is left to future work. Paper Pilot positions LLM-assisted writing as a controlled human-AI decision-support process rather than a fully autonomous authorship pipeline.