TIAO:面向文本摘要的令牌重要性感知策略优化
TIAO: Token Importance-Aware Policy Optimization for Text Summarization
浏览论文内容
中文总结 AI 辅助
针对文本摘要中令牌重要性被忽视的问题,提出令牌重要性感知策略优化(TIAO),通过识别核心令牌并重新加权优势,使7B模型性能媲美GPT-4和GPT-5-nano。
中文摘要 AI 辅助
文本摘要要求模型在压缩内容的同时保持一致性、连贯性等关键质量。大型语言模型(LLM)在此任务上表现出色,并可通过强化学习(RL)进一步提升。然而,现有大多数方法将奖励信号直接应用于无差别的令牌序列,忽略了单个令牌对摘要中词级和句级质量的不同重要性。本文提出令牌重要性感知策略优化(TIAO),一种新颖的强化学习策略,明确利用令牌重要性感知。具体而言,TIAO基于令牌依赖关系识别核心令牌,并根据其整体依赖关系重新加权轨迹的优势。在真实世界数据集上的实验表明,我们的TIAO取得了极具竞争力的结果,且经TIAO增强的7B基础模型性能可与GPT-4和GPT-5-nano相媲美。代码可在以下网址获取:https://this URL
英文摘要
Text summarization requires models to condense content while preserving key qualities such as consistency and coherence. Large language models (LLMs) have shown strong performance on this task and can be further improved through reinforcement learning (RL). However, most existing methods apply reward signals directly to undifferentiated token sequences, overlooking the varying importance of individual tokens to word and sentence level quality in summarization. In this paper, we propose Token Importance-Aware Policy Optimization (TIAO), a novel reinforcement learning strategy that explicitly leverages token-importance awareness. Specifically, TIAO identifies core tokens based on token dependency and reweights a trajectory's advantage according to its overall dependencies. Experiments on the real world dataset show that our TIAO achieves highly competitive results, and that a 7B foundation model enhanced by TIAO performs comparably to GPT-4 and GPT-5-nano. Code is available at https://github.com/TechCloud-x/TIAO
发表机构
- National University of Defense Technology(国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。