发表机构
Cornell Tech(康奈尔科技学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出将LLM的推测解码模块改造为序列分类器,在可忽略开销下实现高效分类,其小型探测模型在多任务上优于零样本GPT-5.4-mini,部分任务性能匹配专用8B安全分类器。
AI 中文摘要
语言模型推理过程中的实时分类对安全过滤、行为分析和模型监控具有重要价值,但现有方法需在准确率与效率间权衡。隐藏态探测速度快但存在局限:要么不具备上下文感知能力(仅作用于单个向量,无法建模不同位置间的交互),要么成本极高(需使用专用分类器模型,如Llama Guard、Qwen Guard、LLM-as-judge,或对所有token的隐藏态执行计算后再池化结果,如MultiMax),这体现了效率与准确率间的固有权衡。然而,研究发现近期大语言模型中的推测解码模块可被重新用于高效高质量分类:通过在目标序列末尾附加训练好的软提示,可将推测解码模块改造为序列分类器;在推测解码流水线的推理阶段,键值缓存已存储在GPU内存中,因此分类仅产生可忽略的开销。研究在四项分类任务、四个模型(Qwen3.5-4B、9B、27B、MiniCPM4.1-8B)上进行评估,小型探测模型的表现始终优于零样本GPT-5.4-mini,且在多语言提示安全任务上,无需运行完整大语言模型即可达到或超越专用8B安全分类器(Qwen3Guard-Gen-8B、Llama-Guard-3-8B)的性能。
英文摘要
Real-time classification during language model inference is valuable for safety filtering, behavioral analysis, and model monitoring, but current approaches force a trade-off between accuracy and efficiency. Hidden-state probes are fast but limited: they are either not context-aware: operating on a single vector and cannot model interactions across positions; or they are very costly: having dedicated classifier models (Llama Guard, Qwen Guard, LLM-as-judge) or performing computation on hidden states for all tokens and then pooling the results (MultiMax). This shows an intrinsic trade-off between efficiency and accuracy. However, we find that the speculative-decoding module in recent LLMs can be repurposed for efficient high-quality classification. By appending a trained soft prompt at the end of the target sequence, we can repurpose the speculative-decoding module into a sequence classifier. At inference time in a speculative-decoding pipeline, the KV cache is already in GPU memory, so classification adds negligible overhead. We evaluate on four classification tasks across four models (Qwen3.5-4B, 9B, 27B, MiniCPM4.1-8B). Our small probes consistently outperform zero-shot GPT-5.4-mini and, on multilingual prompt safety, match or beat specialized 8B safety classifiers (Qwen3Guard-Gen-8B, Llama-Guard-3-8B) without running a full LLM.