arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过两步验证增强生成式信息抽取:产品属性用例

Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case

Yi-Sheng Hsu, Nermeen Abou Baker, Uwe Handmann

arXiv 2607.26780首次发表:更新:

发表机构

Computer Science Institute, Ruhr West University of Applied Sciences(鲁尔西应用科学大学计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对数字产品护照领域标注数据稀缺问题,提出集成PLM模块的两步验证生成式信息抽取方法,可提升LLM对低显著性实体的抽取性能,中型模型表现接近大型模型,最小型开源模型效果有限,还开发了对应演示应用。

AI 中文摘要

大型语言模型(LLM)处理与生成文本的能力为信息抽取(IE)应用带来了潜力。尽管LLM是否在分类任务中优于小型微调模型尚存争议,但其强大的泛化能力使其在微调可用标注数据有限的领域颇具前景,这一优势对新兴的数字产品护照(DPP)应用尤为重要——该问题空间广泛但领域特定数据稀缺。受此用例驱动,我们将生成式IE应用于产品领域,明确应对效率、泛化性和数据隐私约束。我们提出一种两步验证方法,在生成式IE流程中集成PLM模块,从而利用LLM的校正能力。我们发现,此类验证任务可提升LLM性能,尤其针对文本中稀疏出现的弱表达、低显著性实体的抽取。对于某些实体,中型模型的性能甚至可达到与大型模型相当的水平,且第一步PLM预测的改进也能提升最终LLM输出。不过,该方法对最小型开源LLM(如Llama-3.2 3B)的效果有限。基于这些发现,我们开发了一款利用本地部署LLM的产品信息抽取演示应用,旨在进一步适配真实世界的DPP用例。

英文摘要

The ability of large language models (LLMs) to process and generate text has introduced potential for applications in information extraction (IE). While it's debated whether LLMs outperform smaller fine-tuned models for classification tasks, their strong generalization capability makes them promising for domains with limited labeled data available for fine-tuning. This advantage is particularly relevant for the emerging application of the digital product passport (DPP), where the problem space is broad but domain-specific data remains scarce. Motivated by this use case, we apply generative IE to the product domain, explicitly addressing efficiency, generalizability, and data privacy constraints. We propose a two-step validation method that integrates a PLM block into the generative IE pipeline and thereby leverages LLMs' correction capability. We discover that such a validation task enhances LLM performance, particularly on the extraction of weakly expressed, low-salience entities that appear sparsely throughout the text. For certain entities, the performance of mid-size models can even reach levels comparable to larger models, and the improvement of first-step PLM predictions also enhance the final LLM output. Nevertheless, the effects on the smallest open-source LLMs (e.g., Llama-3.2 3B) is limited. Based on the findings, we develop a demo application for product information extraction that utilizes locally deployed LLMs, targeting further adaptations to real-world DPP use cases.

Comments13 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑