ArguLens:用于自动作文评分及感知标签的反馈生成的开源系统
ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出开源可本地部署的ArguLens系统,将自动作文评分分解为三个解耦组件,在PERSUADE 2.0测试集上取得良好评分效果,附带反馈生成器并以Apache 2.0许可发布。
AI中文摘要:
多数自动作文评分(AES)系统仅输出单一整体分数,缺乏可解释证据,且依赖封闭API,带来数据隐私与成本壁垒。本文提出ArguLens,一个开源、可本地部署的系统,将AES分解为三个解耦组件:基于LoRA在PERSUADE 2.0上微调的话语单元分类器(Qwen2.5-7B-Instruct)、基于31个语言与话语特征的独立于分数等级的LightGBM评分器,以及通过vLLM提供服务、以Qwen2.5-14B-Instruct为骨干的感知标签反馈生成器。Gradio网页UI提供可插拔推理后端,支持单篇与批量作文评分,并可下载单篇作文的详细分析结果。在作文不重叠的PERSUADE 2.0测试集上,logitprobe分类器达到82.6%准确率与0.727宏F1值;在按提示分组的5折交叉验证下,评分器在使用最优话语特征协议时的平均QWK为0.813,消融实验显示,添加黄金话语标注相较于词汇+句法配置,QWK提升了0.055(配对t检验,p=0.010)。这是组件级诊断,而非分类器到评分器的端到端结果。反馈生成器附带结构化评估协议,其人工评估研究留待未来工作。该系统以Apache 2.0许可发布,网址为this https URL。
英文摘要:
Most automated essay scoring (AES) systems output a single holistic score without interpretable evidence and rely on closed APIs that introduce data privacy and cost barriers. We present ArguLens, an opensource, locally deployable system that decomposes AES into three decoupled components: a discourse-move classifier (Qwen2.5-7B-Instruct fine-tuned with LoRA on PERSUADE 2.0), a grade-independent LightGBM scorer over 31 linguistic and discourse features, and a label-aware feedback generator served through vLLM with a Qwen2.5-14BInstruct backbone. A Gradio web UI exposes pluggable inference backends and supports single-essay and batch scoring with downloadable per-essay breakdowns. On an essaydisjoint PERSUADE 2.0 test split, the logitprobe classifier achieves 82.6% accuracy and 0.727 macro-F1; under prompt-grouped 5-fold cross-validation the scorer reaches a mean QWK of 0.813 under an oracle discoursefeature protocol, and an ablation shows that adding gold discourse annotations yields an increment of +0.055 QWK over the lexical+syntactic configuration (paired t-test, p = 0.010). This is a component-level diagnostic rather than an end-to-end classifier-to-scorer result. The feedback generator ships with a structured evaluation protocol; its human-rater study is left to future work. The system is released under Apache 2.0 at https://github.com/wwrwbs/AI_AWE.