arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20250cs.LG

亚3B开源语言模型在8 GB消费级GPU上的零样本作文评分能走多远?

How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?

Nguyen Dung Son, Dang Quang Minh, Nguyen Huu Loi, Truong Viet Vu, Nguyen Thai Anh

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在8 GB消费级GPU上评估亚3B开源模型零样本作文评分,发现规则分解提示优于整体提示,多特质特化修复校准问题,但最佳性能仍低于人类和长度基线,定位为形成性反馈。

中文摘要 AI 辅助

使用大型语言模型进行零样本作文评分通常通过专有API模型展示,然而自动评分最需要的场景,如公立学校在严格隐私规则下评分数千篇作文,往往是将学生写作发送给第三方API不可接受的情况。我们研究当模型必须是亚3B开源模型并完全在本地以FP16运行时,有多少能力得以保留,通过对来自两个家族的四个指令微调模型(Qwen2.5在0.5B/1.5B/3B,SmolLM2在1.7B)在单个8 GB消费级GPU上的所有八个ASAP-AES提示进行受控研究,采用自助置信区间、Holm校正配对检验和关键设计选择的部署现实变体。出现三个发现。(i) 在批量最小-最大聚合下,对于每个模型,规则分解提示优于整体提示(尽管Qwen2.5-3B在一个提示上显著下降),在均值聚合下,两个不相关家族在1.5-1.7B规模上相差在0.01以内。(ii) 将特质分数映射到提示范围对评分者校准敏感:一个模型将特质压缩到狭窄的低带(0-10中的2-4),朴素均值聚合崩溃,而多特质特化的最小-最大归一化修复了它(宏QWK从0.204到0.388),当其统计量在30篇留出作文上冻结时保持在0.03以内。(iii) 在十二种配置中的十一种中,带符号误差随作文长度下降,与大型LLM评分者报告的冗长偏差相反;归一化规则分解在很大程度上为校准良好的模型平坦化了这一斜率。我们诚实地锚定结果:最佳本地配置(0.388)仍远低于人类评分者间上限(0.769)和仅长度基线(0.523),因此我们将亚3B本地模型严格定位为形成性、人工监督的反馈。

英文摘要

Zero-shot essay scoring with large language models is usually demonstrated with proprietary API models, yet the settings where automated scoring is most needed, such as public schools grading thousands of essays under strict privacy rules, are often those where sending student writing to a third-party API is unacceptable. We ask how much capability survives when the model must be a sub-3B open model running fully locally in FP16, with a controlled study of four instruction-tuned models from two families (Qwen2.5 at 0.5B/1.5B/3B, SmolLM2 at 1.7B) on all eight ASAP-AES prompts on a single 8 GB consumer GPU, with bootstrap confidence intervals, Holm-corrected paired tests, and deployment-realistic variants of the key design choices. Three findings emerge. (i) Rubric-decomposed prompting beats holistic prompting for every model under batch min-max aggregation (though Qwen2.5-3B drops significantly on one prompt), and under mean aggregation two unrelated families land within 0.01 at the 1.5-1.7B scale. (ii) Mapping trait scores into the prompt range is fragile to grader calibration: one model compresses traits into a narrow low band (2-4 on 0-10) and naive mean aggregation collapses, while the min-max normalization of Multi-Trait Specialization repairs it (macro QWK 0.204 to 0.388) and stays within 0.03 when its statistics are frozen on 30 held-out essays. (iii) Signed error falls with essay length in eleven of twelve configurations, opposite to the verbosity bias reported for large LLM judges; normalized rubric decomposition largely flattens this slope for well-calibrated models. We anchor results honestly: the best local configuration (0.388) remains far below both the human inter-rater ceiling (0.769) and a length-only baseline (0.523), so we position sub-3B local models strictly for formative, human-supervised feedback.

发表机构

  • FPT University(FPT大学)
  • Van Lang University(Van Lang大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑