arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33720cs.CLcs.AI

从细粒度修订操作到有意义的修订单元:评估大语言模型在修订边界检测中的表现

From Granular Revision Operations to Meaningful Revision Units: Evaluating LLMs for Revision Boundary Detection

Yu Tian, Andrew Potter, Katerina Christhilf, Motahareh Darvishpour Ahandani, Jessica Early, Steve Graham, Danielle S. McNamara

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估LLMs在检测写作修订边界中的表现,发现微调Qwen3-32B(宏F1=.859)优于提示方法及邻近启发式,表明任务适应对有效利用LLMs至关重要。

中文摘要 AI 辅助

修订痕迹为学生的写作过程提供了有价值的证据,但其对学习分析的有用性取决于单个修订如何被表示。自动草稿比较方法通常会产生细粒度的编辑操作,这些操作可能将一个有目的的修订分割成多个分析单元。本研究评估了大语言模型(LLMs)是否能在结构化的修订操作数据中识别有意义的修订单元边界,以及它们是否能提供超越简单非LLM基线的价值。使用来自本科生写作的113对匹配的草稿-修订对,专家标注产生了4,344个候选边界。我们比较了零样本和少样本的GPT-5.5与基础Qwen3-32B、Qwen3-32B的参数高效微调,以及基于多数和邻近度的基线。尽管接收了修订上下文和任务指令,没有提示的LLM条件优于邻近启发式(宏F1=.825)。相比之下,使用两种上下文表示的微调Qwen3-32B实现了最高的宏F1(.859),识别了更多的同单元关系,同时保持了与启发式相当的精确度。确定性后处理显著改善了提示模型,但对最强的微调模型几乎没有增加益处。这些发现表明,当任务适应时,LLMs可以支持修订边界判断,但仅靠通用提示可能无法超越透明的结构启发式。

英文摘要

Revision traces provide valuable evidence about students' writing processes, but their usefulness for learning analytics depends on how individual revisions are represented. Automated draft-comparison methods often produce granular edit operations that can fragment a single purposeful revision into multiple analytic units. This study evaluates whether LLMs can identify meaningful revision unit boundaries in structured revision operation data and whether they provide value beyond simple non-LLM baselines. Using 113 matched draft--revision pairs from undergraduate writing, expert annotation yielded 4,344 candidate boundaries. We compared zero- and few-shot GPT-5.5 and base Qwen3-32B, parameter-efficient fine-tuning of Qwen3-32B, and majority and proximity-based baselines. Despite receiving revision context and task instructions, no prompted LLM condition outperformed the proximity heuristic (macro-F1 = .825). In contrast, fine-tuned Qwen3-32B using the two context representation achieved the highest macro-F1 (.859), identifying more same-unit relationships while maintaining precision comparable to the heuristic. Deterministic post-processing substantially improved the prompted models but added little benefit to the strongest fine-tuned model. These findings suggest that LLMs can support revision boundary judgment when task-adapted, but general purpose prompting alone may not outperform transparent structural heuristics.

发表机构

  • Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑