arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Spark-to-Paper:作为可组合技能的端到端研究论文生成系统

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

Zhuoyang Qian, Biao Wu, Yiran Wang, Chris D Yan, Desan Dai, Liangwei Zheng, Jin Jiang, Junsheng Zhang, Wenhao Wang

arXiv 2608.11924首次发表:更新:

发表机构

Vast Intelligence Lab; University of Technology Sydney(星智实验室; 悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Spark-to-Paper是在现有编码助手中实现的端到端研究论文生成系统,含13项可组合技能,在8个研究主题中实现高引用与图表有效性,提升伪造检测率,成本与耗时较低。

AI 中文摘要

将研究创意转化为完整论文,仅靠文本生成是不够的:系统必须检索文献、设计并执行实验、根据证据修正论点、生成可用于发表的图表,还要在漫长的生成过程中保持一致性。我们提出Spark-to-Paper,这是一个端到端研究论文生成系统,它作为13项可组合技能在现有编码助手中实现,无需单独的智能体平台或编排服务。Spark-to-Paper将基于模型的判断与可直接执行和检查的确定性操作分离开来,还将实验规划与报告分离开来,以便在观察结果前指定所需证据,并根据测量结果修正手稿论点。为提升长研究轨迹的可靠性,该系统将确定性完整性检查与自我批判相结合,并限制一种称为“自反驳循环”的故障模式,即重复实验持续拒绝原始研究目标。Spark-to-Paper还通过编程绘图生成实验结果的可编辑矢量图,通过基于代码的重构生成方法图。在8个受控研究主题中,Spark-to-Paper实现了99.5%的引用有效性和96.4%的图表可编辑性;受控消融实验显示,完整的完整性与审查堆栈将伪造检测率从单次草稿的14%提升至92%,而对抗性审查的精度达到74%;完整系统使用1190万token,每份手稿成本8.1美元,平均耗时3.2小时。这些结果表明,端到端研究论文生成可作为轻量、可组合的工作流在现有编码助手中实现,同时将实验证据作为论点被接受、修正或放弃的核心依据。

英文摘要

Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generation process. We present Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service. Spark-to-Paper separates model-based judgment from deterministic operations that can be directly executed and checked. It further separates experiment planning from reporting, so that required evidence is specified before results are observed and manuscript claims are revised according to measured outcomes. To improve reliability over long research trajectories, the system combines deterministic integrity checks with self-critique and bounds a failure mode we call the Self-Refutation Loop, in which repeated experiments continue to reject the original research objective. Spark-to-Paper also produces editable vector figures through programmatic plotting for experimental results and code-based reconstruction for generated method diagrams. Across eight controlled research topics, Spark-to-Paper achieves 99.5% citation validity and 96.4% figure editability. A controlled ablation increases fabrication detection from 14% for a single-pass draft to 92% with the full integrity and review stack, while adversarial review achieves 74% precision. The full system uses 11.9M tokens, costs $8.1 per manuscript, and requires 3.2 hours on average. These results show that end-to-end research paper generation can be implemented as a lightweight, composable workflow inside existing coding assistants while keeping experimental evidence central to how claims are accepted, revised, or abandoned.

Comments24 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑