AI 中文总结
该研究针对LLM生成论文引言的挑战,提出StructPO框架将多阶段写作工作流内化为单遍策略,实验显示其在语义对齐等指标上优于基线,泛化性好且与GPT-5.1性能相当。
AI 中文摘要
利用大语言模型(LLMs)生成严谨的论文引言仍具挑战性,因为它需要在连贯的叙述中协调背景、缺口识别、方法与贡献。现有解决方案将该过程外部化为多阶段提示或智能体工作流,成本高昂且易受跨阶段漂移影响。我们提出StructPO,这是一种结构感知策略学习框架,它将整个多阶段写作工作流内化为由显式阶段标记控制的单遍策略。StructPO引入结构感知信用分配以解耦局部阶段质量与全局连贯性,以及细化引导优化以将修订行为内化为首遍策略。实验表明,StructPO相较于基于工作流的基线,在语义对齐、结构合理性和推理效率上均有提升,能泛化到域外出设置,且当扩展至Qwen3-32B时,在人工评估中仍与GPT-5.1具有竞争力。这些结果表明,通过细粒度策略优化内化学术写作工作流,为成本高昂的外部编排提供了可行替代方案。
英文摘要
Generating a rigorous paper introduction with large language models (LLMs) remains challenging, since it requires coordinating background, gap identification, method and contribution within a coherent narrative. Existing solutions externalize this process as multi-stage prompts or agent workflows which are expensive and vulnerable to cross-stage drift. We propose StructPO, a struct-aware policy learning framework that internalizes the entire multi-stage writing workflow into a single-pass policy controlled by explicit stage tokens. StructPO introduces struct-aware credit assignment to decouple local stage quality from global coherence and refinement-guided optimization to internalize revision behavior into the first-pass policy. Experiments show that StructPO improves semantic alignment, structural rationality and inference efficiency over workflow-based baselines, generalizes to out-of-domain settings, and remains competitive with GPT-5.1 in human evaluation when scaled to Qwen3-32B. These results show that internalizing academic writing workflows through fine-grained policy optimization offers a viable alternative to costly external orchestration.