arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00014cs.AI

CoT-Core:通过感知思维链的核心集选择加速大语言模型评估

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

Qihua Pan, Zhenheng Tang, Peijie Dong, Xiang Liu, Huacan Wang, Bo Li, Xiaowen Chu

首次发表
浏览论文内容

中文总结 AI 辅助

CoT-Core是无训练的核心问题选择框架,通过感知大语言模型的思维链推理轨迹聚类问题,可大幅降低LLMs评估成本,且在GSM8K等多数据集上维持高保真度分数估计。

中文摘要 AI 辅助

大语言模型(LLMs)在持续开发过程中的评估会产生极高的计算开销。尽管核心集选择可加速评估,但现有方法要么存在严重的“冷启动”瓶颈,需要大量历史日志(如项目反应理论),要么表现出表面词汇偏差,忽略了任务的底层推理流形。我们提出CoT-Core,一种新型的无训练核心问题选择框架。鉴于词汇差异大的问题可共享等价的底层逻辑,CoT-Core提示LLMs展开零样本思维链(CoT)推理轨迹,将这些路径投影到潜在空间,从而按内在逻辑等价性而非表面文本相似性对问题进行聚类。在GSM8K、MMLU、MMLU-Pro和GPQA上的大量实验表明,CoT-Core在保持高保真度分数估计的同时大幅降低了评估成本,并明确了感知推理剪枝的边界条件,揭示其效能本质上受任务复杂度的制约。

英文摘要

Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes. While coreset selection accelerates evaluation, existing methods either suffer from a severe ``cold start'' bottleneck requiring massive historical logs (e.g., Item Response Theory) or exhibit a surface lexical bias that misses the underlying reasoning manifold of tasks. We propose CoT-Core, a novel training-free core question selection framework. Recognizing that lexically disparate questions can share equivalent underlying logic, CoT-Core prompts LLMs to unroll zero-shot Chain-of-Thought (CoT) reasoning trajectories. Projecting these paths into a latent space effectively clusters questions by intrinsic logical equivalence rather than superficial text similarity. Extensive experiments on GSM8K, MMLU, MMLU-Pro, and GPQA demonstrate that CoT-Core drastically reduces evaluation costs while maintaining high-fidelity score estimation, and delineate the boundary conditions of reasoning-aware pruning, revealing that its efficacy is intrinsically gated by task complexity.

发表机构

  • Nanjing University(南京大学)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑