arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提取论点,而非仅分类:指令微调的大语言模型用于生成式组件检测

Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection

Sofiane Elguendouze, Erwan Hain, Elena Cabrio, Serena Villata

arXiv 2609.24855首次发表:更新:

发表机构

CNRS; INRIA; I3S(法国国家科学研究中心; 法国国家信息与自动化研究所; I3S实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ITFACD,利用指令微调的大语言模型将论元组件检测重构为生成任务,直接从文本识别论元,在标准基准上超越现有系统,展示了生成式方法的潜力。

AI 中文摘要

论元组件检测(ACD)是论元挖掘(AM)的核心子任务,也是其最具挑战性的方面之一,因为它需要同时界定论元跨度并将其分类为诸如主张和前提等组件。尽管与其他AM任务相比,针对该子任务的研究仍然相对有限,但大多数现有方法将其表述为简化的序列标注问题、组件分类问题,或先进行组件分割再进行分类的流水线。在本文中,我们提出了ITFACD,一种基于指令微调的大语言模型(LLMs)的新方法,使用紧凑的基于指令的提示,并将ACD重新构建为语言生成任务,使得能够直接从纯文本中识别论元,而无需依赖预分割的组件。在标准基准上的实验表明,我们的方法相比最先进的系统取得了更高的性能。据我们所知,这是首次完全将ACD建模为生成式任务的尝试之一,凸显了指令微调在复杂AM问题上的潜力。我们的代码和使用的数据集在以下GitHub仓库中公开可用。

英文摘要

Argumentative component detection (ACD) is a core subtask of Argument(ation) Mining (AM) and one of its most challenging aspects, as it requires jointly delimiting argumentative spans and classifying them into components such as claims and premises. While research on this subtask remains relatively limited compared to other AM tasks, most existing approaches formulate it as a simplified sequence labeling problem, component classification, or a pipeline of component segmentation followed by classification. In this paper, we propose ITFACD, a novel approach based on instruction-tuned Large Language Models (LLMs) using compact instruction-based prompts, and reframe ACD as a language generation task, enabling arguments to be identified directly from plain text without relying on pre-segmented components. Experiments on standard benchmarks show that our approach achieves higher performance compared to state-of-the-art systems. To the best of our knowledge, this is one of the first attempts to fully model ACD as a generative task, highlighting the potential of instruction tuning for complex AM problems. Our code and the datasets used are openly available in the following GitHub repository.

Journal refCOLM 2026 - Third Annual Conference on Language Modeling, Oct 2026, San Francisco, CA, United States

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑