SyntaxBench:大型语言模型字符级推理的统计诊断框架
SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models
- University of Nevada, Las Vegas(内华达大学拉斯维加斯分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对字符级推理评估不足,提出SyntaxBench诊断基准与统计框架,含六项任务和八模型评估,揭示分词与推理模式的影响,并指出子串提取任务仍具挑战。
AI中文摘要:
大型语言模型越来越多地被用于那些细微语法错误至关重要的场景,然而字符级推理目前仍主要通过孤立的探针和总体准确率来评估。我们提出了SyntaxBench,一个用于字符级推理的诊断基准和统计评估框架。它包含五个核心任务:字符计数、字母包含、回文检测、编辑距离和最长字符串选择,以及index_to_span——一个更难的子串提取压力测试。五个核心任务使用配对的英文输入和字符长度匹配的随机字符串输入。index_to_span文档共享200-500词的词数范围,且不进行字符长度匹配。所有六个任务均使用零样本、单样本和四样本提示。我们评估了从2B到32B参数的八个开放权重模型,涵盖11种推理模式配置。该框架报告精确匹配和宽松准确率、Cohen's kappa、带优势比的配对McNemar检验、自助置信区间、Kendall's tau、类别条件指标、分词分析以及多重比较校正检验。三个发现尤为突出。首先,分词影响准确率:随机字符串比英文字符串更具字符可见性(每词元1.892个字符对比3.169个字符),且随着英文单词占据更多词元,字符计数准确率下降。其次,推理模式并非普遍有益:在近乎饱和的任务上,Gemma4-31B在不同模式下几乎不变,而Qwen3.6-27B在回文检测中启用思考时表现更差(四样本下非思考0.952对比思考0.886)。第三,index_to_span在很大程度上仍未解决;最佳四样本精确匹配准确率为6.75%。字符级评估需要受控输入、配对检验以及对分词和推理模式的分析,而非仅依赖总体准确率。
英文摘要:
Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts. We evaluate eight open-weight models from 2B to 32B parameters across 11 reasoning-mode configurations. The framework reports exact-match and relaxed accuracy, Cohen's kappa, paired McNemar tests with odds ratios, bootstrap confidence intervals, Kendall's tau, class-conditional metrics, tokenization analysis, and multiple-comparison-corrected tests. Three findings stand out. First, tokenization shapes accuracy: random strings are more character-visible than English strings (1.892 vs. 3.169 characters per token), and character-counting accuracy falls as English words occupy more tokens. Second, reasoning mode is not uniformly helpful: Gemma4-31B is nearly unchanged across modes on the near-saturated tasks, while Qwen3.6-27B is worse with thinking on palindrome detection (0.952 non-thinking vs. 0.886 thinking at four-shot). Third, index_to_span remains largely unsolved; the best four-shot exact-match accuracy is 6.75%. Character-level evaluation needs controlled inputs, paired tests, and analyses of tokenization and reasoning mode rather than aggregate accuracy alone.