跳转至

衡量部分得分差距:越南2025年凸性评分方案的严格基准测试

文章背景与核心概要

当大语言模型在人类考试中进行评估时,传统的基准测试通常将回答简单地判定为“正确”或“错误”,并报告一个整体准确率。这种方法依赖于一个假设,即部分知识可以转化为成比例的分数,然而当考试采用非加性(凸性)评分方案时,这一假设便不再成立。

本文深入探讨了越南2025年国家高中毕业考试的改革,旨在揭示标准准确率指标的缺陷。作者引入了 THPT-Ladder 基准测试,涵盖了来自11个学科的21场官方考试中的632个题目,并使用教育部实际应用的评分标准对大语言模型进行了评估。研究结果表明,标准指标通过奖励官方评分标准中明确惩罚的部分知识,显著夸大了模型的得分,从而扭曲了模型的真实能力及其在考生群体中的排名。


📌 摘要

当语言模型在人类考试中进行评估时,传统的基准测试通常将回答简单地判定为“正确”或“错误”,并报告一个简单的整体准确率。这种方法依赖于一个假设,即部分知识可以转化为成比例的分数——当考试实施非加性(凸性)评分方案时,这一假设便不再成立。

When language models are evaluated on human exams, traditional benchmarks typically score responses as strictly correct or incorrect, reporting a simple overall accuracy. This approach relies on the assumption that partial knowledge translates into proportional credit—an assumption that breaks down when examinations implement non-additive (convex) grading schemes.

本文探讨了越南2025年国家高中毕业考试的改革,以突出标准准确率指标的缺陷。通过引入 THPT-Ladder(一个包含11个学科、21场官方考试、632个题目的基准测试),作者使用教育部实际应用的评分标准对语言模型进行了评估。研究发现,标准指标通过奖励官方评分标准中明确惩罚的部分知识,显著夸大了模型的得分,最终扭曲了模型的真实能力和群体排名。

This paper investigates the 2025 reform of Vietnam's National High School Graduation Examination to highlight the flaws of standard accuracy metrics. By introducing THPT-Ladder—a benchmark of 632 items from 21 official exams across 11 subjects—the authors evaluate language models using the exact grading rubrics applied by the ministry. The findings reveal that standard metrics significantly inflate model scores by rewarding partial knowledge that the official rubric actively penalizes, ultimately distorting a model's true competence and cohort ranking.


🔍 关键见解与方法论

  • 凸性评分方案: 在2025年考试的第二部分,考生需对每道题的四个是非判断题进行评估。其评分标准呈凸性,根据正确陈述的数量,分别给予 0、0.10、0.25、0.50 或 1.00 分。例如,答对三个陈述可获得 \(0.50\) 分,而不是标准比例评分下的 \(0.75\) 分。
  • The Convex Marking Scheme: In Part II of the 2025 exam, candidates evaluate four true/false statements per question. The grading scale is convex, awarding 0, 0.10, 0.25, 0.50, or 1.00 points depending on the number of correct statements. For instance, getting three statements right yields \(0.50\) points rather than the \(0.75\) points awarded by standard proportional credit.
  • 对模型评分的影响: 在考试总分 \(10.00\) 分中,该部分占 \(4.00\) 分。与标准比例评分相比,官方评分标准在八个受测模型中,每道第二部分题目平均少给 \(0.020\)\(0.159\)
  • Impact on Model Scoring: Accounting for \(4.00\) out of the exam's \(10.00\) total points, the official rubric pays \(0.020\) to \(0.159\) points less per Part II question across eight tested models compared to standard proportional credit.
  • 扭曲的群体排名: 由于数百万考生的成绩已公开,模型可以直接与人类群体进行基准对比。例如,Qwen3.5-27B 在2025年历史考试中因 \(0.042\) 分的差距,其在 \(481,293\) 名考生中的排名从 第90百分位下降至第77百分位
  • Distorted Cohort Rankings: Because millions of human candidate scores are published, models can be directly benchmarked against human cohorts. For example, a \(0.042\)-point shortfall for Qwen3.5-27B on the 2025 History exam drops its standing from the 90th to the 77th percentile among \(481,293\) candidates.
  • 误差分布敏感性: 模型的整体准确率并不能预测其受到的惩罚程度。在 Claude Sonnet 5 的准确率水平下,不同的错误分布会导致每题得分在 \(0.869\)\(0.932\) 分之间波动,这证明了官方得分在很大程度上取决于正确陈述是如何组合的。
  • Error Distribution Sensitivity: A model's overall accuracy does not predict its penalty. At Claude Sonnet 5's accuracy level, varying error distributions result in scores ranging from \(0.869\) to \(0.932\) points per question, proving that official marks heavily depend on how correct statements are grouped.

🔗 链接与资源