arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00218cs.CL

少量神经元可揭示大语言模型何时滥用工具:用于可靠工具使用的稀疏检测与选择性引导

A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use

Yutong Ke, Ming Yin, Chongwen Zhao, Kaizhu Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出PRISMS框架,利用少量特定MLP神经元实现大语言模型工具使用故障的稀疏检测与选择性引导,可跨模型系列降低过度调用率、提升工具所需准确率。

中文摘要 AI 辅助

具现化智能体大语言模型(Agentic LLMs)会出现三类关键的工具使用故障:无效参数(有效性问题)、不必要的调用(过度调用),以及需要工具时却省略调用(遗漏调用)。我们发现,一小部分针对特定故障的多层感知机(MLP)神经元可通过线性可分决策边界区分此类故障。基于这一观察,我们提出PRISMS(Probing Representations In Support of Monitoring and Steering,用于监测与引导的探测表征),这是一种闭环框架,在稀疏检测与激活引导之间共享特定故障的神经元基础。PRISMS会选择对贡献至关重要的MLP神经元,并对其激活拟合L1正则化检测器。在来自Qwen3、Llama和Gemma系列的6个模型上,过度调用与遗漏调用可在生成前提示边界处被检测,ROC-AUC为0.90至1.00;而有效性可从生成的工具调用跨度中被检测,ROC-AUC为0.86至0.90。这些结果是通过高度稀疏的读数实现的:遗漏调用仅需1-2个MLP神经元,过度调用需2-16个,有效性约需128个。这些稀疏检测器使用的特征比密集残差流基线少23至627倍,且性能相当或更优。共享的神经元基础还支持对工具调用行为的双向控制,抑制不必要的调用并触发遗漏的调用。因此,PRISMS会依据预测的故障风险触发干预,以缓解无条件引导的附带影响。在所有6个模型上,PRISMS将合并的过度调用率降低了80%(从0.131降至0.026),同时将工具所需准确率提高了14.2个百分点(从0.689升至0.831)。PRISMS因此为跨模型系列提供了轻量的故障检测与选择性干预方案。

英文摘要

Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools are needed (missing). We find that a small, failure-specific set of MLP neurons could distinguish such failures with linearly separable decision boundaries. Building on this observation, we introduce PRISMS (Probing Representations In Support of Monitoring and Steering), a closed-loop framework that shares a failure-specific neuron basis between sparse detection and activation steering. PRISMS selects contribution-critical MLP neurons and fits an L1-regularized detector on their activations. Across six models from the Qwen3, Llama, and Gemma families, over-calling and missing are detected at the pre-generation prompt boundary with ROC-AUC 0.90-1.00, while validity is detected from the generated tool-call span with ROC-AUC 0.86-0.90. These results are achieved with highly sparse readouts: only 1-2 MLP neurons for missing, 2-16 for over-calling, and approximately 128 for validity. These sparse detectors match or outperform dense residual-stream baselines using 23-627 times fewer features. The shared neuron basis also supports bidirectional control over tool-calling behavior, suppressing unnecessary calls and eliciting omitted ones. PRISMS therefore gates intervention on predicted failure risk to mitigate the collateral effects of unconditional steering. Across all six models, PRISMS reduces pooled over-calling rate by 80% (from 0.131 to 0.026) while increasing tool-required accuracy by 14.2 percentage points (from 0.689 to 0.831). PRISMS thus provides lightweight failure detection and selective intervention across model families.

补充信息

↑