AI 中文总结
本研究通过实证发现,词元级困惑度的简单片段级聚合与人类代码可理解性相关性不可靠,明确其仅适用于局部代码困惑,需经特定分析才能用于可靠的代码可理解性测量。
AI 中文摘要
近期研究表明,大型语言模型的词元级困惑度与代码理解过程中人类的局部困惑程度相符,这引发了一个自然问题:困惑度是否也能作为代码片段级的可理解性信号?我们针对该问题开展实证研究,涉及多个基于人类标注的数据集,包括方法级可理解性判断、输出预测任务以及已被接受的可理解性改进补丁。尽管已有词元级证据,我们发现词元困惑度的简单片段级聚合方式(如平均、中位数或峰值困惑度)与人类可理解性的相关性并不可靠。随后我们探究了原因:其一,词元困惑度在代码结构中呈高度偏态且为重尾分布,极端峰值不仅源于语义有意义的结构,还来自标识符、字面量、类型、分隔符及分词伪影;其二,人类可理解性标注常缺乏共识,导致整个片段的难度是一个有噪声的目标;其三,困惑度分布及其与人类难度的对齐情况在不同模型和分词器间存在显著差异。这些发现解释了为何先前的词元级困惑度-困惑度对齐无法直接迁移至片段级可理解性。总体而言,本研究将困惑度定位为一种有前景但需谨慎使用的认知信号:它对局部代码困惑有用,但需结合代码感知的聚合方式、共识感知的评估以及模型敏感性分析,才能支持可靠的代码可理解性测量。
英文摘要
Recent work suggests that token-level perplexity from large language models can align with localized human confusion during code comprehension. This raises a natural question: can perplexity also serve as a snippet-level signal for code understandability? We conduct an empirical study of this question across multiple human-grounded datasets, including method-level understandability judgments, output-prediction tasks, and accepted understandability-improvement patches. Despite prior token-level evidence, we find that simple snippet-level aggregations of token perplexity, such as average, median, or peak perplexity, do not reliably correlate with human understandability. We then investigate why this happens. First, token perplexity is highly skewed and heavy-tailed across code structures; extreme spikes arise not only from semantically meaningful constructs, but also from identifiers, literals, types, separators, and tokenization artifacts. Second, human understandability labels often lack consensus, making whole-snippet difficulty a noisy target. Third, perplexity distributions and their alignment with human difficulty vary substantially across models and tokenizers. These findings explain why prior token-level perplexity--confusion alignment does not directly transfer to snippet-level understandability. Overall, our study positions perplexity as a promising but delicate cognitive signal: useful for localized code confusion, but requiring code-aware aggregation, consensus-aware evaluation, and model-sensitivity analysis before it can support reliable code-understandability measurement.