arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习从用户生成的主题语料库中结构化数据

Learning to structure data from user-generated thematic corpora

Elishay Avram, Oren Glickman, Elad Yom-Tov

arXiv 2610.01463首次发表:更新:

发表机构

Bar-Ilan University(巴伊兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出全自动迭代框架,利用大语言模型从主题语料库中无需预定义本体地发现并提取领域属性模式,在Reddit健康社区上属性发现与人类一致性达61%,支持经济高效地构建结构化数据集。

AI 中文摘要

主题语料库(如社交媒体社区)包含描述本可以结构化的数据的非结构化文本。这些数据包括,例如,社交媒体数据中提到的个人属性、行为和经历。提取结构化数据具有挑战性,因为相关属性通常是隐含的、依赖于领域的,并且事先未知。我们提出了一个全自动的、迭代的框架,用于在没有预定义本体的情况下发现和提取特定领域的属性模式。利用大型语言模型(LLMs),该框架归纳候选属性,顺序整合语义重叠的属性,并分配结构类型。这些使得能够创建本体并利用语料库中的值填充它。该框架还使得能够使用较小的LLMs进行值提取,与大型LLMs相比,精度损失可估计。我们在5个与健康相关的Reddit社区上评估了该框架。发现的属性与人类识别的属性的一致性达到61%,接近独立标注者之间62%的一致性。在大多数情况下,算法在少于10次迭代内收敛到稳定的属性集。结构类型分配达到82%的准确率,值提取与人工标注相比达到0.8的F1分数。在四个LLM家族中,较小的指令调优模型在模型规模增加时,相对于高容量参考LLM,在提取性能上显示出统计学上显著的改进,支持明智的精度-成本权衡。这些结果表明,与人类识别的属性相当的属性可以被自动发现,从而能够经济且大规模地创建高质量的结构化数据集。通过消除对预定义本体的需求,迭代的模型驱动的模式归纳为挖掘主题语料库提供了实用且可扩展的基础。

英文摘要

Thematic corpora, such as social media communities, contain unstructured text describing data that could be made structured. These include, for example, personal attributes, behaviors, and experiences mentioned in social media data. Extracting structured data is challenging as relevant attributes are often implicit, domain-dependent, and unknown in advance. We propose a fully automated, iterative framework for discovering and extracting domain-specific attribute schemas without a predefined ontology. Using large language models (LLMs), the framework induces candidate attributes, sequentially consolidates semantically overlapping attributes, and assigns a structural type. These enable creating an ontology and populating it with values from the corpus. The framework also enables the use of smaller LLMs for value extraction with estimable accuracy loss compared to large LLMs. We evaluate the framework on 5 health-related Reddit communities. Discovered attributes achieved 61% agreement with human-identified attributes, close to the 62% agreement between independent annotators. In most cases, the algorithm converges to a stable attribute set in fewer than 10 iterations. Structural type assignment achieves 82% accuracy, and value extraction reaches an F1 score of 0.8 compared to human annotations. Across four LLM families, smaller instruction-tuned models show statistically significant improvements in extraction performance with model scale when evaluated against a high-capacity reference LLM, supporting informed accuracy-cost trade-offs. These results show that attributes comparable to those identified by humans can be discovered automatically, enabling the creation of high-quality structured datasets economically and at scale. By removing the need for predefined ontologies, iterative model-driven schema induction offers a practical and scalable foundation for mining thematic corpora.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑