CoT-Core:通过感知思维链的核心集选择加速大语言模型评估
CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection
浏览论文内容
中文总结 AI 辅助
CoT-Core是无训练的核心问题选择框架,通过感知大语言模型的思维链推理轨迹聚类问题,可大幅降低LLMs评估成本,且在GSM8K等多数据集上维持高保真度分数估计。
中文摘要 AI 辅助
大语言模型(LLMs)在持续开发过程中的评估会产生极高的计算开销。尽管核心集选择可加速评估,但现有方法要么存在严重的“冷启动”瓶颈,需要大量历史日志(如项目反应理论),要么表现出表面词汇偏差,忽略了任务的底层推理流形。我们提出CoT-Core,一种新型的无训练核心问题选择框架。鉴于词汇差异大的问题可共享等价的底层逻辑,CoT-Core提示LLMs展开零样本思维链(CoT)推理轨迹,将这些路径投影到潜在空间,从而按内在逻辑等价性而非表面文本相似性对问题进行聚类。在GSM8K、MMLU、MMLU-Pro和GPQA上的大量实验表明,CoT-Core在保持高保真度分数估计的同时大幅降低了评估成本,并明确了感知推理剪枝的边界条件,揭示其效能本质上受任务复杂度的制约。
英文摘要
Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes. While coreset selection accelerates evaluation, existing methods either suffer from a severe ``cold start'' bottleneck requiring massive historical logs (e.g., Item Response Theory) or exhibit a surface lexical bias that misses the underlying reasoning manifold of tasks. We propose CoT-Core, a novel training-free core question selection framework. Recognizing that lexically disparate questions can share equivalent underlying logic, CoT-Core prompts LLMs to unroll zero-shot Chain-of-Thought (CoT) reasoning trajectories. Projecting these paths into a latent space effectively clusters questions by intrinsic logical equivalence rather than superficial text similarity. Extensive experiments on GSM8K, MMLU, MMLU-Pro, and GPQA demonstrate that CoT-Core drastically reduces evaluation costs while maintaining high-fidelity score estimation, and delineate the boundary conditions of reasoning-aware pruning, revealing that its efficacy is intrinsically gated by task complexity.
发表机构
- Nanjing University(南京大学)
- The Hong Kong University of Science and Technology(香港科技大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。