AI 中文总结
RepBench构建多基准支撑的大语言模型能力表征探测数据层,揭示读出方式与聚合准则对评估的重要性,相关资源以闭环工作流形式发布。
AI 中文摘要
表征工程可读取并引导大语言模型的能力方向,但现有方法通常仅在特定领域的合成数据上评估,导致测量结果难以比较或复现,且可能反映表面模式而非实际能力。本文提出RepBench,这是一个基于基准的能力对齐表征探测数据层:从13427篇基准论文中爬取得到13个家族共182个能力簇的分类体系;从353个公开基准数据集中提取出46149个经审核的探测文本,覆盖94项能力,每项能力均由至少两个独立基准支持。这种多基准设计降低了对单一数据源的依赖:原始文本向量无自然簇粒度,而基准池化后的能力向量在所有12个评估模型上均呈现少量簇的内部聚类最优性,且与人类分类体系一致性较低。在跨基准迁移评估中,12个模型均完成全部4种读出方式,均值差法在10个模型上达到最高模型级均值,逻辑回归则在最多的能力-模型单元中胜出。这种分歧表明,读出方式与聚合准则是重要的评估维度。该流程、语料及评估代码已作为可复用的闭环工作流发布。
英文摘要
Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data. The resulting measurements are difficult to compare or reproduce and may reflect surface patterns rather than capabilities. We present RepBench, a benchmark-grounded data layer for capability-aligned representation probing. Crawling 13,427 benchmark papers yields a taxonomy of 182 capability clusters in 13 families; harvesting 353 public benchmark datasets yields 46,149 audited probe texts covering 94 capabilities, each supported by at least two independent benchmarks. This multi-benchmark design reduces dependence on any single source: raw per-text vectors exhibit no natural cluster granularity, whereas benchmark-pooled capability vectors show an interior clustering optimum at a small number of clusters on all 12 evaluated models, with low agreement to the human taxonomy. Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells. This disagreement shows that the readout method and aggregation criterion are meaningful evaluation dimensions. The pipeline, corpus, and evaluation code are released as a reusable closed-loop workflow.
Comments22 pages, 8 figures, with appendices. Yanshi Li and Xueru Bai contributed equally