PoolBench:面向仅解码器大语言模型概念表示评估的池化策略基准
PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs
浏览论文内容
中文总结 AI 辅助
PoolBench是针对仅解码器大语言模型的基准,固定评估协议隔离池化为变量,对比19种策略等,发现W4_hierarchical性能优于基线,构建方法影响更大,发布相关资源。
中文摘要 AI 辅助
池化是仅解码器模型概念表示研究中一个重要却未被充分研究的设计选择:从业者必须将令牌级隐藏状态折叠为段落级向量,但目前尚无跨概念、模型和任务比较该选择的统一协议。报告的性能提升会因数据集、层数、构建方法和池化规则的同时变化而混淆,导致无法做出有原则的决策。我们提出PoolBench,一个将池化作为实验变量、在固定评估协议下开展研究的基准。PoolBench涵盖17个概念、19种池化策略和3个开源权重仅解码器模型(Llama-3.1-8B、Gemma-2-9B、Mistral-7B),在单一经过审核的包含37693段真实文本的语料库上进行评估。主要评估维度为线性可分性(D1/AUROC);引导概念流行度(D2/SCP)和输出级解耦(D3)作为诊断维度。核心发现明确:W4_hierarchical的跨模型平均AUROC达到0.7799,而广泛采用的基线P1_last_token仅为0.7640,且存在统计学显著差异(Friedman+Nemenyi检验,p=2.0e-36;18种有效策略中有77对显著差异)。排名在各层间保持稳定(rho=0.961--0.990)。一项关键负面结果:强检测性能不代表强引导性能——对大多数概念而言,D2和D3的表现远弱于D1,这表明是根本的表示限制而非池化失败。在中等难度概念上,W4_hierarchical较P1_last_token的AUROC提升0.042--0.113;构建方法选择(DiffMean vs. REPE)的影响(AUROC差值0.15)大于池化的影响(AUROC差值0.016),确立了正确的实践优先级。我们发布该语料库、预提取激活、评分模型、引导向量和评估代码,作为池化研究的可复用协议。
英文摘要
Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks. Reported gains are confounded by simultaneous changes in dataset, layer, construction method, and pooling rule, making principled decisions impossible. We introduce PoolBench, a benchmark that isolates pooling as the experimental variable under a fixed evaluation protocol. PoolBench covers 17 concepts, 19 pooling strategies, and 3 open-weight decoder-only models (Llama-3.1-8B, Gemma-2-9B, Mistral-7B), evaluated on a single audited corpus of 37,693 real-text passages. The primary axis is linear separability (D1/AUROC); steered concept prevalence (D2/SCP) and output-level disentanglement (D3) serve as diagnostic axes. The primary finding is decisive: W4_hierarchical reaches a cross-model mean AUROC of 0.7799, while the widely adopted P1_last_token baseline reaches only 0.7640 and is statistically significantly worse (Friedman+Nemenyi, p = 2.0e-36; 77 significant pairs among 18 effective strategies). Rankings are stable across layers (rho = 0.961--0.990). A key negative result: strong detection does not imply strong steering -- D2 and D3 are substantially weaker than D1 for most concepts, indicating a fundamental representational limit rather than a pooling failure. On mid-difficulty concepts, W4_hierarchical outperforms P1_last_token by 0.042--0.113 AUROC; construction method choice (DiffMean vs. REPE) has a larger effect (delta AUROC 0.15) than pooling (delta AUROC 0.016), establishing the correct practical hierarchy. We release the corpus, pre-extracted activations, scorer models, steering vectors, and evaluation code as a reusable protocol for pooling research.