STAR-OPD:面向ABSA四元组抽取的结构化级联感知在线策略奖励蒸馏
STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction
- Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对ABSA四元组抽取中蒸馏模型的结构错误问题,提出STAR-OPD方法,通过在线策略蒸馏结合级联感知集结构奖励,在多个数据集上优于基线,缩小了学生-教师差距并提升了推理效率。
AI中文摘要:
基于方面的情感分析(ABSA)四元组抽取需要在包含多个细粒度情感元组的评论上联合预测目标、方面、观点和情感。尽管大型思维链(CoT)模型在该任务上表现良好,但将其蒸馏为更小的可部署模型仍存在困难。我们发现了蒸馏ABSA抽取中特定于任务的失败模式:学生模型在目标-方面接口处的错误会产生结构无效状态,例如断裂的目标-方面绑定和幻觉目标,进而破坏下游预测。传统的离线策略蒸馏不适合该场景,因为它仅在教师生成的轨迹上训练,且对主导推理的学生诱导结构状态提供的监督很少。为解决这种不匹配,我们提出STAR-OPD(STructured Aspect-cascade-aware On-Policy Reward Distillation,结构化级联感知在线策略奖励蒸馏),它基于通用在线策略蒸馏构建,并为ABSA四元组抽取实例化为级联感知、集结构奖励。STAR-OPD在学生模型的 rollout(展开)上训练,并应用直接针对绑定一致性、目标 grounding( grounding 指接地,即与真实信息对齐)和细粒度方面消歧的集结构奖励。在E-ABSA20K和SemEval-2014上的实验表明,STAR-OPD始终优于离线策略和通用在线策略基线,减少了目标幻觉,并显著提高了结构困难案例的性能。使用Qwen3-4B时,STAR-OPD大幅缩小了学生-教师差距,同时提高了推理效率,凸显了在线策略结构校正对蒸馏ABSA抽取的重要性。
英文摘要:
Aspect-based sentiment analysis (ABSA) quadruple extraction requires jointly predicting target, aspect, opinion, and sentiment over reviews that often contain multiple fine-grained sentiment tuples. While large chain-of-thought (CoT) models perform well on this task, distilling them into smaller deployable models remains difficult. We identify a task-specific failure mode in distilled ABSA extraction: student errors at the target-aspect interface create structurally invalid states, such as broken target-aspect bindings and hallucinated targets, which then corrupt downstream predictions. Conventional off-policy distillation is poorly suited to this setting because it trains only on teacher-generated trajectories and provides little supervision on the student-induced structural states that dominate inference. To address this mismatch, we propose STAR-OPD (STructured Aspect-cascade-aware On-Policy Reward Distillation), which builds on generic on-policy distillation and instantiates it for ABSA quadruple extraction with cascade-aware, set-structured rewards. STAR-OPD trains on student rollouts and applies set-structured rewards that directly target binding consistency, target grounding, and fine-grained aspect disambiguation. Experiments on E-ABSA20K and SemEval-2014 show that STAR-OPD consistently outperforms off-policy and general on-policy baselines, reduces target hallucination, and substantially improves performance on structurally hard cases. With Qwen3-4B, STAR-OPD substantially narrows the student-teacher gap while improving inference efficiency, highlighting the importance of on-policy structural correction for distilled ABSA extraction.