arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06425cs.CLcs.LG

NTDH:用于综合情感分析的复杂推理

NTDH: Complex Reasoning for Comprehensive Affective Analysis

Tianlei Zhu, Zhiwei Liu, Yuyan Wang, Xiao-Yang Liu, Sophia Ananiadou

首次发表
浏览论文内容

中文总结 AI 辅助

针对综合情感分析的挑战,提出NTDH方法,通过自然化、感知容差门等策略,在少数据量下提升性能,获最优EI-reg结果。

中文摘要 AI 辅助

综合情感分析面临两大挑战:其一,它涵盖了具有连续、有序和多标签输出的异质性预测任务;其二,情感含义依赖于上下文,需要协调相互冲突的线索,而非直接映射到标签。现有方法直接学习这种映射,未显式建模协调过程。我们将该任务重构为复杂推理问题,该方法可在异质性标签空间上生成一个输出接口,并提供一条可优化可验证奖励的轨迹;据我们所知,这是首个同时涵盖情感和情绪的此类处理方案。数据方面存在障碍:必须合成情感推理轨迹,而通用合成方式与情感的目标、容差及现象不匹配,且会丢弃或泄露其失败案例。我们提出NTDH,以解决这四类失败问题:自然化将训练答案设为黄金标签,确保其构造层面的正确性;感知容差门按任务自身评分容差检查每个答案;领域感知策略利用情感科学的理念优化推理;方向提示仅报告错误的类型和方向,不暴露目标。我们使用SFT(监督微调)训练Qwen3-8B,随后在验证所用的相同容差下进行GRPO(生成式强化学习策略优化)训练(多标签子任务采用更宽松的构造门),并通过组件消融量化各部分的数据质量影响。使用16302条训练记录(约为同类指令微调系统的1/14),最终策略在6项官方测试指标中的5项上优于其SFT检查点,且在对比系统中取得最强的EI-reg结果,皮尔逊相关系数达0.862。

英文摘要

Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is context-dependent, requiring conflicting cues to be reconciled rather than mapped directly to labels. Existing methods learn this mapping directly and do not model the reconciliation explicitly. We recast the task as a complex-reasoning problem, which yields one output interface across heterogeneous label spaces and a trajectory over which a verifiable reward can be optimised; to our knowledge, this is the first such treatment covering both sentiment and emotion. The obstacle is on the data side: affective reasoning traces must be synthesised, and generic synthesis is misaligned with the targets, tolerances, and phenomena of affect, and discards or leaks its failure cases. We propose NTDH, which addresses these four failures. Naturalisation sets the training answer to the gold label, so it is correct by construction. A Tolerance-aware gate checks each answer against the task's own scoring margin. Domain-aware strategies refine the reasoning using ideas from affective science. Directional Hints report only the type and direction of an error, without exposing the target. We train Qwen3-8B with SFT and then GRPO under the same tolerance used for verification (up to a more permissive construction gate on the multi-label subtask), and a component ablation quantifies the data-quality effect of each part. Using 16,302 training records, about 14x fewer than comparable instruction-tuned systems, the final policy improves over its SFT checkpoint on five of six official-test metrics and achieves the strongest EI-reg result among the compared systems, at a Pearson correlation of 0.862.

发表机构

  • Columbia University(哥伦比亚大学)
  • The University of Manchester(曼彻斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑