TenderKG:基于2021-2023年法国公共采购数据的大规模知识图谱数据集
TenderKG
AI总结:
本文推出TenderKG,这是2021-2023年法国公共采购数据构建的大规模知识图谱数据集,可用于高风险决策场景下的知识感知推荐等研究,为相关方法评估提供基准。
AI中文摘要:
公共采购是一项重要的经济活动,公共机构通过竞争性招标流程将合同分配给企业。尽管该领域十分重要,但推荐系统对其的探索仍较少,主要原因是缺乏能体现其复杂性的公开可用数据集。本文中,我们推出TenderKG,这是一个由2021至2023年法国公共采购数据构建的大规模知识图谱数据集。该数据集通过异构实体对采购生态系统进行建模,包括企业、招标、标段以及工作领域的特定领域分类法,这些实体通过丰富的语义和结构关系相连。该场景的一个关键特点是仅能看到中标企业,导致中标交互的显式信号较为稀疏。为克服这一局限,TenderKG整合了法国招标市场参与者和招标的大量辅助信息,包括文本描述、分层分类和地理特征,从而支持在高度受限和竞争激烈的环境中开展知识感知推荐研究。我们提供了该数据集的详细统计数据和分析,突出了其结构特性、稀疏性模式和特定领域特征。我们认为TenderKG为投标者推荐、基于知识图谱的推荐、竞争感知匹配等方向开辟了新的研究路径,并为评估现实世界高风险决策场景中的方法提供了有价值的基准。
英文摘要:
Public procurement represents a major economic activity, where public institutions allocate contracts to companies through competitive tendering processes. Despite its importance, this domain remains underexplored by recommender systems, largely due to the lack of publicly available datasets capturing its complexity. In this paper, we introduce TenderKG, a large-scale knowledge graph dataset constructed from French public procurement data covering the period 2021--2023. The dataset models the procurement ecosystem through heterogeneous entities, including companies, tenders, lots, and domain-specific taxonomies of work domains, connected via rich semantic and structural relations. A key specificity of this setting is that only the awarded companies are visible, resulting in sparse explicit signals of awarded interactions. To overcome this limitation, TenderKG integrates extensive side information on the actors in the French tender market and the tenders, including textual descriptions, hierarchical classifications, and geographical features, enabling the study of knowledge-aware recommendation in a highly constrained and competitive environment. We provide detailed statistics and analyses of the dataset, highlighting its structural properties, sparsity patterns, and domain-specific characteristics. We believe TenderKG opens new research directions in bidder recommendation, knowledge graph-based recommendation, competition-aware matching, and provides a valuable benchmark for evaluating methods in real-world, high-stakes decision-making scenarios.