arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29584cs.LGcs.CL

一种基于纵向文本数据建模组织级语义身份的计算框架

A Computational Framework for Modelling Organisation-Level Semantic Identity from Longitudinal Textual Data

  • Cardiff University(卡迪夫大学)
  • National Defence University(国防大学)

机构由 AI 辅助整理,请以论文原文为准。

Brinda Murali Krishna, Oktay Karakuş, Can Eyupoglu

AI总结:

该文提出一种整合语义表示学习、图建模与时间演变的计算框架,从纵向文本中推断可解释的组织级语义身份,并以K-pop歌词语料验证其有效性。

AI中文摘要:

组织持续产生大量文本数据,这些数据捕捉了它们如何随时间进行沟通、演变和差异化。尽管自然语言处理的最新进展已大幅提升了组织级文本分析水平,但现有方法主要将组织表示为潜在嵌入或预测性特征向量,用于相似性估计、分类或检索。因此,目前尚无通用的计算框架,可将组织级语义身份建模为一种可解释且不断演变的语义构造,并源自纵向文本证据。本文提出了一种计算框架,在统一的分析方法论中整合了语义表示学习、基于图的语义建模、组织级语义指纹、时间语义演变和证据驱动验证。组织通过描述多样性、集中性、连通性、新颖性和语义社区组成的互补语义维度进行刻画,并对其纵向分析以推断有证据支持的语义身份。该框架使用来自韩国四大娱乐公司旗下艺术家的K-pop歌词纵向语料库进行了演示。实证分析揭示了可区分的多维语义身份、多样的时间演变轨迹和连贯的综合身份概况。全面验证表明,推断出的身份在统计上得到支持,在替代分析假设下具有稳健性,可复现且具有操作信息性。除案例研究外,所提出的框架确立了组织级语义身份,并提供了一种可迁移的方法论,用于从纵向文本数据建模组织行为。

英文摘要:

Organisations continuously generate large volumes of textual data that capture how they communicate, evolve and differentiate themselves over time. Although recent advances in natural language processing have substantially improved organisation-level text analytics, existing approaches primarily represent organisations as latent embeddings or predictive feature vectors for similarity estimation, classification or retrieval. Consequently, there is currently no general computational framework for modelling organisation-level semantic identity as an interpretable and evolving semantic construct derived from longitudinal textual evidence. This paper introduces a computational framework that integrates semantic representation learning, graph-based semantic modelling, organisation-level semantic fingerprints, temporal semantic evolution and evidence-driven validation within a unified analytical methodology. Organisations are characterised through complementary semantic dimensions describing diversity, concentration, connectivity, novelty and semantic community composition, which are analysed longitudinally to infer evidence-supported semantic identities. The framework is demonstrated using a longitudinal corpus of K-pop lyrics from artists affiliated with the four major South Korean entertainment companies. The empirical analyses reveal distinguishable multidimensional semantic identities, diverse temporal evolutionary trajectories and coherent integrated identity profiles. Comprehensive validation demonstrates that the inferred identities are statistically supported, robust under alternative analytical assumptions, reproducible and operationally informative. Beyond the case study, the proposed framework establishes organisation-level semantic identity and provides a transferable methodology for modelling organisational behaviour from longitudinal textual data.

↑