arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SWIM:基于熟练度条件生成的学生写作模拟

SWIM: Student Writing Simulation via Proficiency-Conditioned Generation

Heejin Do, Jakub Kontak, Mrinmaya Sachan

arXiv 2609.03215首次发表:更新:

发表机构

ETH Zurich; ETH AI Center(苏黎世联邦理工学院; ETH人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出SWIM任务,评估提示、SFT、RL方法在学生写作模拟中的表现,发现明确监督比单独提示对齐效果更强,但低熟练度写作仍难复现。

AI 中文摘要

写作熟练度体现在学生构建内容、组织思路、选词及使用语言的方式中。尽管基于大语言模型(LLM)的学生模拟研究日益受到关注,但LLM能否在长篇写作中复现这种多维度差异仍未得到充分探索。本研究探讨语言模型能否真实模拟学生写作,并提出SWIM任务,将学生写作模拟(Student Writing sIMulation)表述为基于熟练度条件的文章生成。我们使用自动作文评分作为画像对齐的衡量标准,评估提示、监督微调(SFT)和强化学习(RL)方法在写作模拟中的表现。大量实验表明,即使是基于评分标准策略的强大专有LLM,提示也仅能提供有限的熟练度控制,具体而言,模型虽能调整内容导向的特征,但难以复现不同熟练度水平下的词汇、语法及组织差异。监督微调大幅提升了对齐效果,而采用所提出的熟练度对齐奖励的强化学习,在所有写作特征及文章提示上均取得了进一步提升。我们的研究结果表明,与单独使用提示相比,明确的监督能实现更强的画像对齐,但真实的低熟练度写作仍难以复现。

英文摘要

Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplored. In this work, we explore if language models can realistically simulate student writing, and introduce SWIM, a task that formulates Student Writing sIMulation as proficiency-conditioned essay generation. We evaluate prompting, supervised fine-tuning (SFT), and reinforcement learning (RL) methods for writing simulation using automated essay scoring as a measure of profile alignment. Extensive experiments reveal that prompting provides limited proficiency control, even for strong proprietary LLMs with rubric-grounded strategies. In particular, while models can adjust content-oriented traits, they struggle to reproduce the lexical, grammatical, and organizational variation in different proficiency levels. SFT substantially improves alignment, while RL with the proposed proficiency-alignment reward yields further gains across all writing traits and essay prompts. Our findings suggest that explicit supervision enables substantially stronger profile alignment than prompting alone, while authentic low-proficiency writing remains challenging to reproduce.

CommentsEMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑