恢复智能体主权:通过对比认知解码缓解共识悖论
Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding
浏览论文内容
中文总结 AI 辅助
针对LLM易受群体共识影响的谄媚问题,提出对比认知解码(CED)零样本干预,通过双前向传播抑制有毒共识,在多个基准上显著提升性能并恢复智能体主权。
中文摘要 AI 辅助
大型语言模型(LLMs)表现出对对抗性群体共识的参数化脆弱性。为缓解这种谄媚行为,我们引入了对比认知解码(CED),一种零样本推理干预方法。与依赖较弱辅助模型的标准对比解码(CD)不同,CED利用单一架构上的双前向传播来隔离从众偏差。通过引入新颖的非对称、零边界概率钳制和离散top-k截断掩码,CED在数学上抑制了有毒共识令牌而不会导致语法崩溃。在复杂基准(GAIA、SWE-bench、Multi-Challenge)上使用Gemma-2(9B)、Llama-3.1(8B)和Mistral v0.3(7B)对7,200对轨迹进行评估,CED成功中和了架构和位置偏差。通过将认知懈怠绝对降低高达33.00%,CED带来了显著的性能提升,实现了高达30.75%的准确率恢复。恢复主权引发了不同的架构行为——Gemma-2中的被动任务聚焦和Llama-3.1中对模拟群体的主动反驳——表明CED无需微调即可将合规性与能力解耦。
英文摘要
Large language models (LLMs) exhibit a parametric vulnerability to adversarial swarm consensus. To mitigate this sycophancy, we introduce Contrastive Epistemic Decoding (CED), a zero-shot inference intervention. Unlike standard Contrastive Decoding (CD) which relies on a weaker secondary model, CED utilizes a dual forward-pass on a single architecture to isolate conformity bias. By introducing a novel asymmetric, zero-bounded probability clamp and discrete top-k truncation mask, CED mathematically suppresses toxic consensus tokens without causing grammatical collapse. Evaluated across 7,200 paired trajectories on complex benchmarks (GAIA, SWE-bench, Multi-Challenge) using Gemma-2 (9B), Llama-3.1 (8B), and Mistral v0.3 (7B), CED successfully neutralizes architectural and positional biases. By reducing cognitive loafing by up to 33.00% absolute, CED drives significant performance gains, yielding up to a 30.75% accuracy recovery. Regaining sovereignty induces distinct architectural behaviors---passive task-focus in Gemma-2 and active refutation of the simulated swarm in Llama-3.1---showing CED decouples compliance from capability without fine-tuning.
发表机构
- University of Waterloo(滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。