发表机构
Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LentEx是首个从大语言模型视角处理潜在实体抽取的框架,通过合成数据与指令微调优化高效LLM,在MTEB聚类基准上超越现有最优模型,泛化性强,可应用于RAG等NLP任务。
AI 中文摘要
潜在实体抽取(LEE)旨在识别自由文本中隐含的、需通过上下文推断的实体,这是传统实体抽取方法难以胜任的领域。本文提出LentEx,一种用于潜在实体抽取的新型框架,利用合成数据生成与指令微调来优化规模较小、高效的大语言模型(LLM)。潜在实体通常抽象且具有主题性,对检索增强生成(RAG)、客户画像分析、知识图谱丰富等应用至关重要。LentEx通过基于模板的方法生成多样、上下文丰富的合成数据,解决了标注数据集稀缺的问题,确保数据具有高变异性并贴合真实世界分布。据我们所知,LentEx是首个从LLM视角系统处理LEE问题的方法。LentEx在多项任务中展现出显著的性能提升,尤其在MTEB聚类基准上超越了现有最优模型。此外,该方法对未见过的领域具备强泛化能力,使LentEx能高效应用于RAG、聚类等真实世界NLP任务,从而为自然语言处理中的潜在实体理解与抽取建立了新范式。
英文摘要
Latent entity extraction (LEE) tackles the challenge of identifying implicit, contextually inferred entities within free text-an area where traditional entity extraction methods fall short. In this paper, we introduce LentEx, a novel framework for latent entity extraction that leverages synthetic data generation and instruction fine-tuning to optimize smaller, efficient large language models (LLMs). Latent entities, which are often abstract and thematic, are crucial for applications such as retrieval-augmented generation (RAG), customer persona analysis, and knowledge graph enrichment. LentEx addresses the scarcity of labeled datasets by employing a template-based approach to generate diverse, contextually rich synthetic data, ensuring high variability and alignment with real-world distributions. To our knowledge, LentEx is the first to systematically approach LEE through the lens of LLMs. LentEx demonstrates significant performance improvements across multiple tasks, notably surpassing state-of-the-art models on the MTEB Clustering Benchmark. Furthermore, our methodology enables robust generalization to unseen domains, making LentEx highly applicable in real-world NLP tasks, including RAG and clustering, thereby establishing a new paradigm for latent entity understanding and extraction in natural language processing.
CommentsPublished in IJCNN 2025. ©2025 IEEE
Journal ref2025 International Joint Conference on Neural Networks (IJCNN), pp. 1-8, 2025
DOI:10.1109/IJCNN64981.2025.11228382