角色感知启发式情景注意力用于对话式大语言模型
Role-aware Heuristic Episodic Attention for Conversational LLMs
- National Key Laboratory of Parallel and Distributed Computing(并行与分布式计算国家重点实验室)
- College of Computer Science and Technology(计算机科学与技术学院)
- National University of Defense Technology(国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多轮对话中上下文衰减问题,提出REA框架,通过指令与情景分离存储及启发式检索,在Long-MT-Bench+上评分提升16.5%,延迟降低2.91倍。
AI中文摘要:
随着多轮对话的增长,大语言模型常常会丢失持续的指令和相关信息。我们通过三种相关的失败模式来研究这种累积性上下文衰减:注意力污染、稀释和漂移。我们提出REA(角色感知启发式情景注意力),这是一个上下文管理框架,为指令和情景交互分配不同的持久性和表示策略。指令记忆在专用前缀中保留已识别的全局约束。情景记忆保留用户输入并压缩模型回复,而启发式检索为每个历史轮次选择原始文本、压缩表示或省略。在Long-MT-Bench+上,REA将裁判评分从6.32提升至7.36(10分制),相对于Vanilla基线相对提升16.5%,并将平均延迟降低2.91倍。额外评估显示在三个覆盖1.7B-7B参数的骨干模型以及中英文角色扮演任务上均有总体收益。这些结果支持角色感知上下文管理作为维持对话连续性和指令遵循的实用方法。
英文摘要:
Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attention), a context-management framework that assigns different persistence and representation policies to instructions and episodic interactions. Instructional Memory retains identified global constraints in a dedicated prefix. Episodic Memory preserves user inputs and compresses model replies, while heuristic retrieval selects raw text, compressed representations, or omission for each historical turn. On Long-MT-Bench+, REA improves the judge score from 6.32 to 7.36 on a 10-point scale, a 16.5% relative gain over the Vanilla baseline, and reduces average latency by 2.91$\times$. Additional evaluations show aggregate gains on three backbones spanning 1.7B-7B parameters and on Chinese and English role-playing tasks. These results support role-aware context management as a practical approach to maintaining conversational continuity and instruction adherence.