发表机构
Xi’an Jiaotong University; Peking University; Tencent(西安交通大学; 北京大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型推理失败问题,提出SPARK方法,利用隐藏状态响应诊断推理状态并引导测试时干预,通过长度控制敏感性等手段,在实验中提升了Qwen3系列模型性能,证明敏感性对推理失败诊断及干预的作用。
AI 中文摘要
大语言模型中的推理失败通常从最终答案评估,但错误答案无法揭示失败原因。现有方法多在输出层面操作,通用激活引导方法未诊断哪些示例需干预。本文介绍SPARK,利用隐藏状态响应诊断模型是否进入有效推理状态并引导轻量级测试时引导。原始隐藏状态敏感性受提示长度强烈混淆,SPARK用长度控制敏感性分离输入规模效应与残余推理激活,结合该信号与跨层协调选择推理活跃锚点和未充分激活的难示例。通过实验,该方法持续提升Qwen3系列模型性能,表明敏感性不仅可作为推理失败的诊断信号,还可作为针对性测试时干预的实用指南。
英文摘要
Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may reflect missing capability, an unstable reasoning trajectory, or a failure to activate a reasoning state that is already available in the frozen model. Existing prompting and benchmark-based evaluation methods mostly operate at the output level, while generic activation-steering methods typically apply global directions without diagnosing which examples require intervention. In this paper, we introduce SPARK, which uses hidden-state response to diagnose whether a model internally enters an effective reasoning state and to guide lightweight test-time steering. The key observation is that raw hidden-state susceptibility is strongly confounded by prompt length, especially in programmatic and algorithmic reasoning where harder serialized instances naturally become longer. SPARK therefore uses length-controlled susceptibility to separate input-scale effects from residual reasoning activation, and combines this signal with cross-layer coordination to select reasoning-active anchors and under-activated hard examples. We use FRONTIER-4.5K as a controlled programmatic reasoning suite for latent profiling and difficulty-aware analysis, and evaluate SPARK-Steering on GSM8K and MATH-500 with forward-only benchmark profiling. Our method improves Qwen3 series models consistently; on MATH-500, accuracy rises from 82.0% to 84.6% for Qwen3-4B and from 82.4% to 85.6% for Qwen3-8B. These results suggest that susceptibility can serve not only as a diagnostic signal for reasoning failures, but also as a practical guide for targeted test-time intervention.