arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GEIS:用于长篇文章生成的智能体技能生成-评估-改进循环

GEIS: A Generation-Evaluation-Improvement Loop of Agent Skills for Long-Form Article Generation

Jiale Zhang, Juntao Hu, Zhijian Ou

arXiv 2607.11503首次发表:更新:

发表机构

Speech Processing and Machine Intelligence (SPMI) Lab, Tsinghua University, China; TasiTech Co., Ltd., China(清华大学语音处理与机器智能实验室; 泰赛科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长篇文章生成难题,提出GEIS循环,通过在Tasi Harness中实现文章写作、证据图像收集等技能,经核心写作、成对评估及改进技能,在20个主题实验中提升了PDF质量得分,证明可将其转变为可改进循环。

AI 中文摘要

长篇文章生成对大语言模型来说仍然困难,因为它需要处理长上下文、长指令和长输出。现有的多智能体管道如STORM虽能提高信息覆盖,但能力受限于提示和固定程序。本文提出GEIS,一种用于维基百科风格长篇文章生成的命名和声明性技能循环。在Tasi Harness中实现并评估,GEIS包括文章写作、基于浏览器的证据和图像收集、图表渲染、PDF感知成对评估以及规则级技能改进等技能。其核心写作技能遵循请求、计划、草稿、审核、完善和交付;成对评估技能产生结构化质量报告;改进技能将反复出现的发现映射为对写作技能的永久补丁。在20个维基百科特色文章主题上评估GEIS,结果表明它能将长篇生成从固定工作流程转变为可检查、模块化和评估引导的改进循环。

英文摘要

Long-form article generation remains difficult for large language models because it combines long context, long instructions, and long outputs. Existing multi-agent pipelines such as STORM improve information coverage by simulating role-specialized agents, but their capabilities are often entangled in prompts and fixed procedures, making them hard to inspect, reuse, or iteratively improve. This paper presents GEIS (Generation-Evaluation-Improvement loop of agent Skills), a loop of named and declarative skills for Wikipedia-style long-form article generation. Implemented and evaluated in Tasi Harness, GEIS composes skills for article writing, browser-based evidence and image collection, diagram rendering, PDF-aware pairwise evaluation, and rule-level skill improvement. Its core writing skill follows Request, Plan, Draft, Audit, Refine, and Deliver; the pairwise evaluation skill produces structured quality reports; and the improvement skill maps recurrent findings into permanent patches to the writing skill in our 20-topic experiment. We evaluate GEIS on 20 Wikipedia Featured Article topics. Under the same generation backend, GEIS improves over the Tasi Harness default writer by 8.0 points on a 100-point PDF quality rubric and outperforms STORM on the two comparable writing dimensions, structural quality and content quality. In the 20-topic improvement experiment, the patched writing skill raises the average score from 82.90 to 86.95, with 17 out of 20 topics improved and the gain mainly coming from content quality. These results show that long-form generation can be reframed from a fixed workflow into an inspectable, modular, and evaluation-guided improvement loop.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑