可操作的幻觉检测:将潜在不确定性转化为智能体批判
Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique
浏览论文内容
中文总结 AI 辅助
该研究提出Latent Critic轻量级LoRA适配器,可实时检测LLM智能体的幻觉,定位准确率超80%、AUROC达0.966,能拦截幻觉并助力智能体自我修正,效能优于同类基线方法。
中文摘要 AI 辅助
部署为智能体的大语言模型(LLMs)常出现用户规范落地失败,执行幻觉生成的不合规动作以强行解决问题,而非表达不确定性。现有检测方法要么无法定位幻觉,要么推理延迟过高,无法提供可操作的实时修正。我们提出Latent Critic,一种轻量级低秩适配器(LoRA),与冻结的基础LLM生成同步运行,主动重构Transformer的残差流——放大潜在落地信号并将其转化为单序列内的本地化自然语言反馈。通过优化基础模型的原生不确定性信号,这种潜在空间操作无需额外推理循环即可实现可靠的细粒度检测。激活补丁和逐层探测的机制分析表明,该秩不变行为将现有不确定性几何重构为线性可分表示,其迁移可靠性优于仅使用基础模型表示。以工具调用作为细粒度幻觉的实例,我们在基于Qwen和Llama的模型上验证了Latent Critic架构的检测能力及下游改进效果。该方法表现出优异的实时效能,在隔离幻觉方面显著优于同等规模的微调外部检测器、语义熵基线和被动内部探测器,达到0.966的AUROC及>80%的定位准确率(例如,未落地内容:日期)。当部署在闭环ReAct环境中时,该Critic作为延迟可忽略的防护机制,在执行前拦截幻觉以防止不合规动作,同时利用该特定本地化反馈实现智能体的高效自我修正。
英文摘要
Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty. Existing detection methods fail to provide actionable, real-time correction as they either do not localize the hallucinations, or incur prohibitive inference latency. We introduce the Latent Critic, a lightweight low-rank adapter (LoRA) that operates concurrently with a frozen base LLM's generation to actively restructure the transformer's residual stream---amplifying latent grounding signals and translating them into localized, natural language feedback within a single sequence. By refining the base model's native uncertainty signals, this manipulation of the latent space enables reliable, granular detection without the overhead of secondary inference loops. Mechanistic analysis via activation patching and layer-wise probing shows that this rank-invariant behavior restructures pre-existing uncertainty geometry into a linearly separable representation that transfers more reliably than base model representations alone. Using tool-calling as an instantiation of granular hallucinations, we validate the detection and downstream improvements enabled by the Latent Critic architecture across Qwen and Llama-based models. Demonstrating superior real-time efficacy, our approach significantly outperforms equivalent-scale fine-tuned external detectors, semantic entropy baselines, and passive internal probes in isolating hallucinations, achieving 0.966 AUROC and >80% accuracy in localization (e.g., ungrounded: date). When deployed in a closed-loop ReAct environment, the Critic acts as a negligible latency guardrail, intercepting hallucinations before execution to prevent undesired actions while simultaneously leveraging this specific localized feedback to enable efficient agent self-correction.
发表机构
- Samsung Research America(美国三星研究院)
机构由 AI 辅助整理,请以论文原文为准。