PULSE:利用稀疏自编码器识别示范效用特征
PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders
浏览论文内容
中文总结 AI 辅助
PULSE提出基于稀疏自编码器的框架,通过定位模型内部与示范效用相关的特征来选择示范,在分类、生成和推理任务上超越最强基线,并验证了特征的预测性和部分跨数据集迁移性。
中文摘要 AI 辅助
上下文学习对示范选择高度敏感,然而大多数方法使用外部查询-示范相似性来选择示范。此类标准可能遗漏模型特定的信号:相似的示范可能激活不同的内部特征和下游行为。我们提出了PULSE(基于稀疏编码的配对效用定位),这是一个基于SAE的框架,用于识别与示范效用相关的模型内部特征,并利用这些特征进行示范选择。利用一个小型带标签的发现集,PULSE采样候选示范集,测量它们在目标模型下的零样本相对效用,并根据特征激活差异与效用差异的对齐程度对SAE特征进行评分。得分最高的正负坐标构成一个稀疏的效用定位向量。我们以两种互补的方式使用该向量:作为有符号分数用于受控的完整集排序,以及作为PULSE-Retriever,将其幅度转换为特征相关性掩码,用于可扩展的池级检索。在分类、生成和推理基准上,PULSE-Retriever相比最强基线分别提升了2-3个准确率点、0.6-0.9个BLEU-4分数和3.2个精确匹配点,同时受控排序验证了所识别特征编码了预测性的集级效用信号。特征检查和跨数据集实验表明,所识别的特征捕获了任务相关、数据集条件化的模式,但保留了部分跨数据集转移的效用信号。我们的代码可在以下网址获取:此https URL。
英文摘要
In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query-demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce PULSE (Paired Utility Localization over Sparse Encodings), an SAE-based framework for identifying model-internal features associated with demonstration utility and using them for demonstration selection. Using a small labeled discovery set, PULSE samples candidate demonstration sets, measures their zero-shot-relative utility under the target model, and scores SAE features by how their activation differences align with utility differences. The top positive and negative coordinates form a sparse utility-localization vector. We use this vector in two complementary ways: as a signed score for controlled complete-set ranking, and as PULSE-Retriever, which converts its magnitude into a feature-relevance mask for scalable pool-scale retrieval. Across classification, generation, and reasoning benchmarks, PULSE-Retriever improves over the strongest baseline by 2-3 accuracy points, 0.6-0.9 BLEU-4, and 3.2 exact-match points, respectively, while controlled ranking validates the identified features encode a predictive set-level utility signal. Feature inspection and cross-dataset experiments suggest that the identified features capture task-relevant, dataset-conditioned patterns, yet retain utility signals that partially transfer across datasets. Our code is available at https://github.com/aohenuo/PULSE.