arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于跨语言立场检测的基于理由引导的知识蒸馏

Rationale-Guided Knowledge Distillation for Cross-Lingual Stance Detection

Qiuli Zhou, Jingyuan Yao, Shengeng Tang, Hongzhi Chen, Jun Tang, Richang Hong

arXiv 2607.18693首次发表:更新:

发表机构

School of Foreign Studies, Hefei University of Technology; School of Computer Science and Information Engineering, Hefei University of Technology; Hefei Institute of Technology(合肥工业大学外国语学院; 合肥工业大学计算机与信息工程学院; 合肥学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对跨语言立场检测中现有方法忽视推理过程及大语言模型存在局限性的问题,提出基于理由引导的知识蒸馏框架,用思维链提示引导大语言模型生成理由并提炼知识至学生模型,设计双路径蒸馏机制及对比学习策略,实验证明该方法优于基线。

AI 中文摘要

立场检测旨在识别文本对给定目标的赞成或反对态度,是各种下游应用的重要任务。现有研究在单语言环境中表现出色,但许多低资源语言缺乏训练有效模型的标注数据。跨语言立场检测可缓解此问题,但现有方法多依赖文本与目标的语义对齐,忽视可靠立场推理所需的推理过程。大语言模型虽有强大推理能力,但计算成本高、推理延迟大。为此,我们提出用于跨语言立场检测的基于理由引导的知识蒸馏框架。具体而言,我们使用思维链提示引导大语言模型生成信息性理由,并将推理知识提炼到紧凑的学生模型中。我们还设计了双路径蒸馏机制来对齐基于理由增强和无理由的表示及其预测分布。此外,引入两种对比学习策略来提高立场辨别能力。多语言基准测试实验表明,我们的方法始终优于有竞争力的基线。

英文摘要

Stance detection aims to identify whether a text expresses a favorable or opposing attitude toward a given target, and serves as an important task for various downstream applications. Although existing studies have achieved strong performance in monolingual settings, especially in English, many low-resource languages such as Catalan still lack sufficient annotated data for training effective models. Cross-lingual stance detection alleviates this problem by transferring stance knowledge from resource-rich languages to low-resource languages. However, most existing methods mainly rely on semantic alignment between texts and targets, while ignoring the reasoning process required for reliable stance inference. Although Large Language Models provide strong reasoning ability, their high computational cost and inference latency limit practical deployment. To address these limitations, we propose a rationale-guided knowledge distillation framework for cross-lingual stance detection. Specifically, we use Chain-of-Thought prompting to guide Large Language Models in generating informative rationales, and distill the resulting reasoning knowledge into a compact student model. We further design a dual-path distillation mechanism to align rationale-enhanced and rationale-free representations, together with their prediction distributions. In addition, two contrastive learning strategies are introduced to improve stance discrimination. Experiments on multilingual benchmarks demonstrate that our method consistently outperforms competitive baselines.

Comments23 pages, 7 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑