发表机构
ScaDS.AI, Technische Universität Dresden(ScaDS.AI,德累斯顿工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于大语言模型的科学文本简化,提出人工参与工作流程,先以GPT-4o-mini生成基线简化,再经不同阶段读者与专家反馈编辑,发布语料库及评估结果,支持跨学科科学交流简化系统相关工作。
AI 中文摘要
跨学科研究正在加速,但科学论文在其所属领域之外仍难以理解。我们研究了基于大语言模型(LLM)的科学文本简化,并提出了一种人工参与的工作流程,将专家摘要转化为非专业人士更易理解的版本。以SciSummNet作为源语料库,首先用GPT-4o-mini生成基线简化。在第一阶段,来自计算机科学以外STEM领域的读者识别难句和短语,并比较原文和GPT简化摘要在可理解性、自然度和简洁性方面的情况。第二阶段,计算机科学专家利用这些反馈创建专家编辑的参考简化。我们发布了由此产生的语料库以及人工判断和自动评估结果。第一阶段的判断显示在可理解性和简洁性方面明显偏好GPT生成的摘要,而对第二阶段编辑的定性分析突出了保留特定领域术语和科学主张力度的重要性。所得资源支持跨学科科学交流简化系统的训练和基准测试。
英文摘要
Interdisciplinary research is accelerating, yet scientific papers remain difficult to understand outside their home fields. We study large language model (LLM)-based simplification of scientific texts and present a human-in-the-loop workflow that transforms expert summaries into more accessible versions for non-specialists. Using SciSummNet as the source corpus, we first generate baseline simplifications with GPT-4o-mini. In Phase 1, readers from STEM fields outside computer science identify difficult sentences and phrases and compare the original and GPT-simplified summaries in terms of comprehensibility, naturalness, and simplicity. In Phase 2, computer science experts use this feedback to create expert-edited reference simplifications. We release the resulting corpus together with human judgments and automatic evaluation results. The Phase 1 judgments show a clear preference for the GPT-generated summaries in terms of comprehensibility and simplicity, while qualitative analysis of the Phase 2 edits highlights the importance of preserving domain-specific terminology and the strength of scientific claims. The resulting resource supports the training and benchmarking of simplification systems for cross-disciplinary scientific communication.
CommentsAccepted at FGWM@KI2026