基于大语言模型的长文本生成中大纲阶段的多框架对比
A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs
浏览论文内容
中文总结 AI 辅助
本研究构建统一基准对比7种长文本生成框架在3种粒度下的大纲表现,提出锚定的大语言模型评估协议,发现框架性能与粒度匹配度相关,大纲与写作表现仅中等相关。
中文摘要 AI 辅助
长文本生成暴露出大语言模型的根本性局限:即使是70B参数的模型,在生成16k token的输出时也会出现长度崩溃,多章节故事则频繁触发“中间丢失”效应特有的属性漂移。“先大纲、后写作”的范式已被广泛采用,但现有研究评估的是最终写作内容而非大纲本身,混淆了本应解耦的两个评估对象。我们构建了一个统一的直接对比基准,涵盖7种代表性长文本生成框架,覆盖3种生成粒度——单章节、多章节和整本书,并提出一种基于锚点的大语言模型评估协议,该协议以5分锚定尺度直接评估大纲与源文本的匹配度。在21种框架-粒度组合中,没有单一框架占据主导地位;性能取决于框架的固有输出形式与目标粒度的匹配程度。SuperWriter在长度受限的单章节模式中排名第一,但该优势在整本书模式中会减弱。大纲侧的排名与写作侧的排名仅存在中等程度的相关性,支持大纲-写作解耦原则。计算限制将写作侧评估限制在部分案例中;后续实验将扩大样本量并增加跨模型评估器,以实现更可靠的统计推断。
英文摘要
Long-form generation exposes fundamental limitations of large language models. Even 70B-parameter models exhibit length collapse at 16k-token outputs, and multi-chapter stories frequently trigger the attribute drift characteristic of the ``lost-in-the-middle'' effect. The ``outline-first, write-later'' paradigm has gained wide adoption, yet existing research evaluates the final writing rather than the outline itself, conflating two evaluation objects that should be decoupled. We construct a unified head-to-head benchmark covering 7 representative long-form generation frameworks across 3 generation granularities -- single-chapter, multi-chapter, and whole-book -- and propose an anchor-based LLM-as-a-judge protocol that directly assesses outlines against the source text on a 5-point anchored scale. Across 21 framework-granularity cells, no single framework dominates; performance depends on the match between a framework's intrinsic output form and the target granularity. SuperWriter ranks first in the length-constrained single-chapter mode, but this advantage degrades in whole-book mode. The outline-side ranking correlates only moderately with the writing-side ranking, supporting the outline--writing decoupling principle. Compute constraints limit the writing-side evaluation to a subset of cases; follow-up experiments will expand the sample size and add cross-model evaluators to enable stronger statistical inference.
发表机构
- Taiyuan Institute of Technology(太原工业学院)
机构由 AI 辅助整理,请以论文原文为准。