发表机构
Visa Research(维萨研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM编码行为差异难以用性能指标区分的问题,提出CLIC方法,通过词元频率分析和决策树区分模型,并定义鲁棒性与集中度指标,开发可视化系统,在22个Kaggle任务上比较10个LLM,为模型选择和提示工程提供见解。
AI 中文摘要
大型语言模型(LLM)在编码任务上的评估主要集中于诸如pass@k之类的性能指标。随着LLM的持续进步,许多模型现已达到基线性能要求,这降低了仅基于性能评估的区分能力。然而,一个关键问题在很大程度上仍未得到探索:LLM在编码行为上有何不同?我们提出CLIC(用于识别与比较的代码学习),一种通过词元频率分析来刻画LLM编码行为的可视化分析方法。CLIC将每个代码样本表示为词元频率的特征向量,并训练一棵可解释的决策树来区分两个LLM的代码集。除分类准确性外,我们定义了两个新指标:鲁棒性,衡量当两个LLM最具区分性的词元被逐步移除时,它们是否仍可区分;以及集中度,衡量差异是由少数主导词元驱动还是分散在众多词元中。解释众多成对比较(跨LLM对、任务和词元化级别)并追踪完整分析链,本质上是一项多尺度、假设驱动的探索任务。因此,我们开发了一个交互式可视化分析系统,以导航比较景观、识别感兴趣的对,并深入探究区分性词元及其代码上下文。对22个Kaggle机器学习任务中10个LLM进行比较的案例研究,为LLM选择和提示工程提供了可操作的见解。
英文摘要
The evaluation of large language models (LLMs) on coding tasks has primarily focused on performance metrics such as pass@k. As LLMs continue to advance, many models now meet baseline performance requirements, reducing the discriminative power of performance-based evaluation alone. Yet a key question remains largely unexplored: how do LLMs differ in their coding behavior? We propose CLIC (Code Learning for Identification and Comparison), a visual analytics approach that characterizes LLM coding behavior through token-frequency analysis. CLIC represents each code sample as a feature vector of token frequencies and trains an interpretable decision tree to separate two LLMs' code sets. Beyond classification accuracy, we define two new metrics: robustness, which measures whether the two LLMs remain distinguishable as their most-discriminative tokens are progressively removed, and concentration, which measures whether the difference is driven by a few dominant tokens or spread across many. Interpreting numerous pairwise comparisons (across LLM pairs, tasks, and tokenization levels) and tracing the full analytical chain form an inherently multi-scale, hypothesis-driven exploration task. We therefore develop an interactive visual analytics system to navigate the comparison landscape, identify pairs of interest, and drill down into discriminative tokens and their code contexts. Case studies comparing 10 LLMs across 22 Kaggle ML tasks reveal actionable insights for LLM selection and prompt engineering.
Comments11 pages, 9 figures