arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从人格中解耦逻辑:边缘LLM智能体对上下文污染的结构性免疫

Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution

Masaaki Nakatsu, Reno Wang

arXiv 2610.09772首次发表:更新:

发表机构

AO, Inc.; OrbLabs AG(AO公司; OrbLabs股份公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对边缘LLM智能体在长历史、误导性人格上下文中的逻辑退化问题,提出解耦架构(AO-DA),将逻辑推理与人格表达分离为两条路径,实验证明其在8B模型上对污染具有结构性免疫,并支持快速人格热切换。

AI 中文摘要

边缘设备上的小型语言模型智能体必须同时保持人格并正确推理,且处于一个被对话历史和人格指令填充的上下文窗口内。我们研究了当历史记录冗长、具有误导性且人格密集时,此类智能体的逻辑部分会发生什么(人格-逻辑干扰),并提出了一种解耦架构(AO-DA),该架构将逻辑推理(“什么”)与人格表达(“如何”)分离到同一个INT4基础模型上的两条推理路径中,并配有热插拔LoRA适配器。逻辑路径仅接收核心轮次,并输出可验证的结构化状态(微状态);人格路径则结合完整历史以角色身份渲染该状态。在Apple M2笔记本电脑上(Llama-3.1-8B-Instruct和Gemma-3-4B-it,4位;4个污染级别×3个臂×2个任务×2个人格×5个种子共480次运行)的同一基础模型消融实验中,我们发现:(i)解耦的逻辑路径对污染具有结构不变性:其提示保持为180(Llama)或167(Gemma)个令牌,而混合单遍提示从242增长到1,203个令牌,且其输出在各级别间字节完全相同(40/40);(ii)混合单遍性能单调下降(Llama上复合逻辑得分从0.669降至0.150,Gemma上从0.487降至0.150),主要原因是未能输出所需的结构化结果(Llama上80-95%的运行,Gemma上在两个最高级别时100%的运行);(iii)将相同的污染输入解耦逻辑路径时,专用适配器、专用格式的路径在8B模型上仍比单遍更稳健(失败率0-20%对比80-95%;配对Δ+0.30至+0.50,Cliff's δ 0.50-0.85,Holm校正p≤0.03),但在4B模型上则不然,两者均崩溃。分离在主题首轮额外增加一次解码(Llama上28.2秒对比18.2秒),并实现了1.7毫秒内的人格热切换,而无需重新运行逻辑路径。代码、评分标准、测试夹具、适配器和日志均已发布。

英文摘要

Small language-model agents on edge devices must hold a persona and reason correctly at once, inside one context window that fills with conversational history and persona instructions. We study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters. The logic path receives only the core turn and emits a verifiable structured state (Micro-State); the persona path renders it in character with the full history. In same-base-model ablations on an Apple M2 laptop (Llama-3.1-8B-Instruct and Gemma-3-4B-it, 4-bit; 480 runs over 4 pollution levels x 3 arms x 2 tasks x 2 personas x 5 seeds) we find: (i) the decoupled logic path is structurally invariant to pollution: its prompt stays at 180 (Llama) or 167 (Gemma) tokens while the mixed single-pass prompt grows from 242 to 1,203, and its outputs are byte-identical across levels (40/40); (ii) the mixed single pass degrades monotonically (composite logic score 0.669 to 0.150 on Llama, 0.487 to 0.150 on Gemma), mostly by failing to emit the required structured output (80-95% of runs on Llama, 100% on Gemma at the two highest levels); (iii) with the same pollution fed into the decoupled logic path, the dedicated-adapter, dedicated-format path is still more robust than the single pass on the 8B model (failure 0-20% vs 80-95%; paired $Δ$ +0.30 to +0.50, Cliff's $δ$ 0.50-0.85, Holm-adjusted $p \le 0.03$) but not on the 4B model, where both collapse. Separation costs one extra decode on a topic's first turn (28.2 s vs 18.2 s on Llama) and buys persona hot-swapping in 1.7 ms without re-running the logic path. Code, rubric, fixtures, adapters and logs are released.

Comments28 pages, 3 figures. Experiment code, scoring rubric, pollution fixtures, adapters and run logs are released (see Appendix G)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑