发表机构
University of Macau; South China University of Technology(澳门大学; 华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出CLAIM流水线,通过生成-攻击协议和纠错规则库,在K-12评估题目生成中达到97.8%通过率,但揭示评分依赖评审者且填空生成存在能力分层与开放集缺陷。
AI 中文摘要
我们提出了CLAIM,一个用于K-12评估题目生成的生产流水线,它结合了双阶段“生成-然后-攻击”协议(模型先作为“课程架构师”起草,然后以敌对的对抗性评审者身份重新进入同一对话)、对已接受和已拒绝题目的双向少样本条件设置(后者携带评估者的诊断),以及一个包含44,844条从该反馈中挖掘并针对每个标准和题目类型检索的纠错规则的知识字典。在755个Common Core ELA标准、三种题目类型和十个LLM上的43,227道已评分题目中,该流水线在9,074道题目的生产运行中达到了97.8%的专家评估者通过率。然后我们追问这个比率认证了什么。使用来自其他供应商的三位评审者对分层样本进行重新评分,且对部署的判定不知情,在每个评审者下都复现了格式排序,并恢复了比部署评估者更大的开放集缺陷;但在生产流行率下,对接受/拒绝二元的判定一致性较弱(kappa约为0.13),且评审者之间的一致性也不更好。因此,该水平是相对于评审者的,且由于没有学生响应数据,我们的质量证据全程都是评估者评判的。该语料库还暴露了一个稳健的不对称性。多项选择和多项选择生成在十几个静态规则下,对于两个前沿模型都饱和在98%或以上,而填空生成则是按能力分层的(在匹配的规则集、标准和评审者下,五个模型的准确率在82.8%到96.7%之间),并且在仅提示优化的条件下达到平台期,随着规则积累,错误分布在答案键过度包含和遗漏之间转移。我们将其分析为开放集边界确定,这是自回归解码器在结构上不适合解决的任务,并展示了当评估者本身被蒸馏时这种不对称性会重现:失败召回率从8%上升到63%,而F1饱和在0.25。
英文摘要
We present CLAIM, a production pipeline for K-12 assessment-item generation coupling a two-stage generate-then-attack protocol (the model drafts as a "curriculum architect", then re-enters the same conversation as a hostile adversarial reviewer), bi-directional few-shot conditioning on accepted and rejected items, the latter carrying the evaluator's diagnosis, and a knowledge dictionary of 44,844 error-correction rules mined from that feedback and retrieved per standard and item type. Across 43,227 scored items over 755 Common Core ELA standards, three item types, and ten LLMs, the pipeline reaches a 97.8% expert-evaluator pass rate on a 9,074-item production run. We then ask what that rate certifies. Re-scoring a stratified sample with three judges from other vendors, blind to the deployed verdict, reproduces the format ordering under every judge and recovers a larger open-set deficit than the deployed evaluator does; but agreement on the accept/reject binary is weak at production prevalence (kappa about 0.13), and the judges agree with each other no better. The level is therefore judge-relative, and with no student-response data our quality evidence is evaluator-judged throughout. The corpus also exposes a robust asymmetry. Multiple-choice and multiple-select generation saturate at 98% or above for both frontier models under a dozen static rules, whereas fill-in-the-blank generation is capability-tiered (82.8-96.7% across five models under a matched rule set, standards, and judge) and plateaus under prompt-only optimization, with error mass shifting between answer-key over-inclusion and omission as rules accumulate. We analyze this as open-set boundary determination, a task autoregressive decoders are structurally ill-equipped to solve, and show the asymmetry recurring when the evaluator itself is distilled: fail-recall rises from 8% to 63% while F1 saturates at 0.25.
Comments31 pages, 8 figures, 16 tables. Companion paper: arXiv:2608.01000