测量部分学分差距:越南2025年凸评分制的严格基准
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme
浏览论文内容
中文总结 AI 辅助
该研究针对越南2025年凸评分制,构建THPT-Ladder基准,发现标准准确率指标会夸大模型在该评分制下的表现,不同错误分布会导致模型得分差异显著,影响模型能力评估。
中文摘要 AI 辅助
在对语言模型进行人类考试评估时,基准通常将每个回答判为正确或错误,并报告整体准确率。这种方法假设部分知识值得获得相应比例的学分,但当考试使用非累加评分制时,这一假设不成立。越南2025年全国高中毕业考试的改革体现了这种替代的代价:考试第二部分中,考生需对每个问题的四个判断陈述进行正误判断,评分采用凸分制,即答对的陈述数量对应0、0.10、0.25、0.50或1.00分,答对三个陈述得0.50分,而非标准准确率指标会给出的0.75分。由于第二部分占考试总分10.00分中的4.00分,报告准确率会因奖励国家明确惩罚的部分知识而夸大分数。我们推出THPT-Ladder基准,包含来自11个科目的21份官方考试的632个题目,评分方式完全按照教育部对学生的评分规则进行。教育部公布了超过100万考生的成绩,使我们能够将模型直接置于人类考生队列中。在8个模型中,官方评分规则对第二部分每个问题的学分比比例学分少0.020至0.159分,这一缺口改变了模型的表观能力:Qwen3.5-27B在2025年历史考试中,0.042分的缺口使其在481293名考生中的排名从第90百分位降至第77百分位。模型的准确率无法预测这一惩罚,在Claude Sonnet 5的准确率水平下,不同的错误分布会导致每个问题的得分在0.869至0.932分之间变化。官方评分取决于正确陈述的分组方式,意味着标准基准报告的是机构不会认可的能力。
英文摘要
When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part II accounts for 4.00 of the exam's 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students. The ministry publishes the marks of over a million candidates, allowing us to place models directly into the human cohort. Across eight models, the official rubric pays 0.020 to 0.159 points less per Part II question than proportional credit. This shortfall changes a model's apparent competence. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile among 481,293 candidates. A model's accuracy does not predict this penalty. At Claude Sonnet 5's accuracy level, different distributions of errors yield scores varying from 0.869 to 0.932 points per question. Official marks depend on how correct statements are grouped, meaning standard benchmarks report a competence the institution would not certify.
发表机构
- Posts and Telecommunications Institute of Technology(越南邮电技术学院)
机构由 AI 辅助整理,请以论文原文为准。