arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PA-CDM:用于评估手写数学表达式识别的位置感知字符检测匹配

PA-CDM: Position-Aware Character Detection Matching for Evaluating Handwritten Mathematical Expression Recognition

Shiliang Luo

arXiv 2609.12917首次发表:更新:

发表机构

East China Normal University(华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出位置感知指标PA-CDM,结合字符检测匹配与位置森林编码,解决手写数学表达式识别评估中位置盲区问题,与人类判断相关性最高(rho=0.9535),优于传统指标且零成本。

AI 中文摘要

手写数学表达式识别(HMER)传统上通过精确匹配率和字符串相似度指标进行评分,但这些指标对错误发生的位置不敏感:两个具有相同标记错误数量的预测,无论它们是错放了下标还是交换了分数的操作数,都会获得相同的分数。基于渲染的字符检测匹配(CDM)能够稳健地对齐字形,但仍然缺乏位置感知——在受控的分数操作数交换测试中,其得分为0.8595,而位置感知评分为0.6253。树编辑指标表现出互补的盲点:解析器归一化覆盖范围之外的重写会被惩罚为结构错误(得分为0.8552,而基于渲染的指标得分为1.0)。我们提出了PA-CDM,一种将字符检测匹配与位置森林编码和散度级加权相结合的位置感知指标;StructPerturb v2.0,一个包含15个类型-强度单元中1,340个受控扰动对的冻结基准;以及一个结合敏感性矩阵、人类研究和LLM评判校准的跨指标一致性协议。在一项六名标注者的研究中,PA-CDM在七种自动指标中与人类判断的相关性最高(Spearman rho=0.9535,n=990)。前沿LLM评判的相关性略高(rho=0.9613),但成本高昂、非确定性且依赖API;PA-CDM以零边际成本接近该性能,且行为确定、可诊断。

英文摘要

Handwritten mathematical expression recognition (HMER) is conventionally scored by exact-match rates and string-similarity metrics that are blind to where an error occurs: two predictions with identical token-error counts receive identical scores whether they misplace a subscript or swap the operands of a fraction. Render-based character detection matching (CDM) aligns glyphs robustly but remains position-blind---on controlled fraction-operand swaps it scores 0.8595 where position-aware scoring yields 0.6253. Tree-edit metrics exhibit a complementary blind spot: rewrites outside the parser's normalization coverage are penalized as structural errors (0.8552 where render-based metrics score 1.0). We propose PA-CDM, a position-aware metric that couples character detection matching with position-forest encoding and divergence-level weighting; StructPerturb v2.0, a frozen benchmark of 1,340 controlled perturbation pairs across 15 type--intensity cells; and a cross-metric consistency protocol combining a sensitivity matrix, a human study, and LLM-judge calibration. In a six-annotator study, PA-CDM attains the highest correlation with human judgments among seven automatic metrics (Spearman rho=0.9535, n=990). A frontier LLM judge correlates slightly higher (rho=0.9613) but is costly, nondeterministic, and API-dependent; PA-CDM approaches it at zero marginal cost with deterministic, diagnosable behavior.

Comments8 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑