LA-RL:信息抽取中强化学习的标签感知自反思
LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型信息抽取中现有反思校正方法不足,提出LA-RL框架,通过任务诊断标签指导自校正,经特定训练阶段提升抽取质量,在多项任务实验中取得良好效果,且发现反思结构对任务敏感。
中文摘要 AI 辅助
大语言模型在信息抽取方面展现出强大潜力,但现有基于反思的校正方法常与结构化抽取输出不匹配。自由形式的自反思能标记错误,却难以确定失败原因。我们引入LA-RL(标签感知反思强化学习),这是一个结果监督框架,用基于任务的诊断标签指导信息抽取自校正。单个主干先预测抽取结果,诊断特定任务错误标签,再依诊断修正输出。训练从注释模型标记的诊断数据开始进行冷启动监督微调,经两个GRPO阶段,奖励最终抽取质量、格式有效性和首次通过的正确性,无需过程奖励模型。在命名实体识别、关系抽取和事件抽取上的实验表明,与SFT相比,同一主干有持续提升,如在SciER关系抽取上平均F1提高6.83,在分布外关系抽取上提高约20 F1,在DuEE1.0上触发F1为14.80,论据F1为17.50。消融实验表明反思结构对任务敏感:更强约束利于关系抽取,而命名实体识别在领域转移下需要限制较少的校正。
英文摘要
Large language models show strong promise for information extraction (IE), but existing reflection-based correction methods are often misaligned with structured extraction outputs. Free-form self-reflection can flag an error, yet it rarely identifies whether the failure is a missing span, wrong label, boundary mismatch, invalid relation type, or reversed argument order. We introduce LA-RL (Label-Aware Reflective Reinforcement Learning), an outcome-supervised framework that guides IE self-correction with task-grounded diagnostic labels. A single backbone first predicts an extraction, diagnoses task-specific error labels, and then revises its output conditioned on the diagnosis. Training starts from diagnostic data labeled by an annotation model for cold-start supervised fine-tuning and proceeds through two GRPO stages that reward final extraction quality, format validity, and first-pass correctness, without a process reward model. Experiments on named entity recognition, relation extraction, and event extraction show consistent same-backbone gains over SFT, including 6.83 average F1 on SciER relation extraction, about 20 F1 on out-of-distribution relation extraction, and 14.80 trigger F1 plus 17.50 argument F1 on DuEE1.0. Ablations show that reflection structure is task-sensitive: stronger constraints benefit relation extraction, whereas named entity recognition needs less restrictive correction under domain shift.
发表机构
- Hefei University of Technology(合肥工业大学)
- Chongqing Jiaotong University(重庆交通大学)
机构由 AI 辅助整理,请以论文原文为准。