FirmCORe:一个关于企业间合作机会结构化推理的基准
FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities
浏览论文内容
中文总结 AI 辅助
FirmCORe是一个人工标注基准,含2,805对企业档案,用于评估LLM推理企业间合作机会;最强模型检测机会宏F1为74.51,但具体类型和角色方向预测准确率较低。
中文摘要 AI 辅助
关于企业间关系的全面结构化数据往往稀缺或难以获取,因为许多关系是私下协商、选择性披露的,并且分散在专有数据库中。这种稀缺性阻碍了合作机会的发现,尤其是对初创企业和中小企业而言。企业档案很容易获得,但合作潜力不能仅从业务相似性推断,因为相似的企业可能是竞争对手,而不相似的企业可能提供互补的产品、技术、渠道、能力或资本。我们提出了FirmCORe(企业间合作机会推理),一个针对弱结构化企业档案进行成对推理的人工标注基准,包含2,805个标注的企业对。给定两个企业档案,模型必须确定现有证据是否支持合作机会,对于正例,还需联合预测其强度、主要合作类型和角色方向。FirmCORe还提供了平行的中英文评估集,包含相同的实例和黄金标签,从而能够对输入语言敏感性进行受控分析。使用具有代表性的本地部署和托管大型语言模型(LLM)进行的实验表明,最强的模型在机会检测上达到了74.51的宏F1分数,但在所有四个输出字段上的精确匹配仅为61.57%。语言效应因模型而异,高跨语言一致性可能掩盖跨语言共享的错误。这些结果表明,当前LLM在检测广泛合作机会方面比识别其具体类型和角色方向要可靠得多。
英文摘要
Comprehensive structured data on inter-firm relationships is often scarce or inaccessible because many relationships are privately negotiated, selectively disclosed, and fragmented across proprietary databases. This scarcity hinders the discovery of collaboration opportunities, particularly for startups and small and medium-sized enterprises. Firm profiles are readily available, but collaboration potential cannot be inferred from business similarity alone, since similar firms may be competitors, whereas dissimilar firms may offer complementary products, technologies, channels, capabilities, or capital. We present FirmCORe (Inter-Firm Collaboration Opportunity Reasoning), a human-annotated benchmark for pairwise reasoning over weakly structured firm profiles, comprising 2,805 labeled firm pairs. Given two firm profiles, a model must determine whether the available evidence supports a collaboration opportunity and, for positive pairs, jointly predict its strength, primary collaboration type, and role direction. FirmCORe also provides parallel Chinese- and English-language evaluation sets containing identical instances and gold labels, enabling controlled analysis of input-language sensitivity. Experiments with representative locally deployed and hosted large language models (LLMs) show that the strongest model achieves a macro-F1 score of 74.51 for opportunity detection but only 61.57% exact match across all four output fields. Language effects vary across models, and high cross-language agreement can mask errors shared across languages. These results indicate that current LLMs are substantially more reliable at detecting broad collaboration opportunities than at identifying their specific types and role directions.
发表机构
- Southwest University of Finance and Economics(西南财经大学)
- Nanyang Technological University(南洋理工大学)
- Beijing University of Post and Telecommunication(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。