StrategyBench:评估大语言模型中的显式策略归纳能力
StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出StrategyBench基准,评估大语言模型显式策略归纳能力,通过多维度分析揭示策略效用的影响因素,相关基准已公开。
中文摘要 AI 辅助
随着大语言模型在数据稀缺和动态变化的任务场景中应用日益广泛,少样本上下文学习(ICL)已成为任务适配的关键范式。然而,直接的ICL通常使用少量示例,未显式抽象任务规则,因此对示例构建敏感。相比之下,人类学习者常通过先从示例中总结任务规则再应用于新实例,来降低这种敏感性。为评估该能力,我们提出StrategyBench,从BIG-Bench中选取可归纳策略的任务,构建参考策略,并沿策略质量和下游效用两个维度定义评估指标。我们进一步从任务变化、模型配置和适配设置三个角度分析策略归纳,涵盖类别差异、生成器-执行器选择、演示设计及基于SFT的适配。实验表明,显式策略效用在不同任务类别间差异显著,且取决于策略生成与执行条件。该基准已发布于:this https URL。
英文摘要
As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstracting task rules, making it sensitive to example construction. In contrast, human learners often reduce such sensitivity by first summarizing task rules from examples and then applying them to new instances. To evaluate this ability, we propose StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility. We further analyze strategy induction from three perspectives: task variation, model configuration, and adaptation setting, covering category-wise differences, generator-executor choices, demonstration design, and SFT-based adaptation. Experiments show that explicit strategy utility differs substantially across task categories and depends on both strategy generation and execution conditions. The benchmark is released at: https://anonymous.4open.science/r/StrategyBench-D53C.