arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对抗指令冲突:基于能量的潜在冲突检测

Combating Instruction Conflict via Energy-Driven Latent Conflict Detection

Mingyu Ma, Yuxin Wu, Jingbo Wang, Tianxiao Huang, Leixin Sun, Xiaochuan Shi

arXiv 2609.08646首次发表:更新:

发表机构

Wuhan University; Renmin University of China(武汉大学; 中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM中用户指令覆盖系统约束的冲突问题,提出ELCD响应级潜在冲突检测器,通过拼接最终令牌与均值池化嵌入并优化排序目标,在多种模型上显著提升检测性能。

AI 中文摘要

大型语言模型(LLMs)越来越多地部署在分层指令环境中,但它们仍然容易受到用户指令覆盖系统级约束的冲突影响。现有的防御机制主要侧重于静态输入检查,因此无法检测到响应漂移(Response Drift)现象,即模型的最终响应违反系统级约束,尽管输入看似合规。为弥补这一空白,我们提出了ELCD,一种用于生成后、交付前验证的响应级潜在冲突检测器。给定完整生成输出,ELCD通过拼接最终令牌嵌入与均值池化响应嵌入来构建复合隐藏状态表示,然后优化成对边际排序目标,以在潜在空间中分离合规响应与漂移响应。在从1.5B到14B参数的五种主流LLM上的大量实验表明,ELCD显著优于竞争基线。值得注意的是,它在Llama-2-7B上将PR-AUC提高了约30个百分点,并将Mistral-7B上95% TPR时的假阳性率(FPR95)降至2.67%。这些结果表明,ELCD为开源权重或自托管LLM部署中的潜在指令冲突检测提供了一种有前景的方法。

英文摘要

Large Language Models (LLMs) are increasingly deployed with hierarchical instructions, yet they remain vulnerable to conflicts in which user directives override system-level constraints. Existing defense mechanisms predominantly focus on static input inspection and therefore fail to detect Response Drift, a phenomenon in which the model's final response violates system-level constraints despite seemingly compliant inputs. To bridge this gap, we introduce ELCD, a response-level latent conflict detector for post-generation, pre-delivery verification. Given the full generated output, ELCD constructs a composite hidden-state representation by concatenating the final-token embedding with the mean-pooled response embedding. It then optimizes a pairwise margin ranking objective to separate compliant and drifting responses in latent space. Extensive experiments across five mainstream LLMs ranging from 1.5B to 14B parameters demonstrate that ELCD significantly outperforms competitive baselines. Notably, it improves the PR-AUC on Llama-2-7B by approximately 30 percentage points and reduces the False Positive Rate at 95% TPR (FPR95) on Mistral-7B to 2.67%. These results suggest that ELCD provides a promising approach for latent instruction-conflict detection in open-weight or self-hosted LLM deployments.

Commentsi need to finish the paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑