基于文本描述符预测初创企业退出:计算语言学框架
Predicting Startup Exit from Textual Descriptors - A Computational Linguistics Framework
浏览论文内容
中文总结 AI 辅助
该研究提出计算语言学框架,仅靠文本描述符结合LightGBM等模型,可在无需财务等变量的情况下预测初创企业退出,还引入了量化的炒作得分,为高信息不对称下的风险投资提供预测信号。
中文摘要 AI 辅助
本研究表明,仅靠文本描述符即可预测早期初创企业的成功,其定义为“退出(Exit)”,无需依赖情境、财务或人力资本变量。研究使用由风险投资机构整理的数据集,涵盖20年间的7419家初创企业,通过初创企业叙事映射分离出基于文本的框架变量,并构建850个特征。对数据子集和向量嵌入进行统计显著性评估后,开展了涵盖6种模型的监督机器学习实验。LightGBM模型实现了最高的预测性能(F1值为0.48),而仅使用文本描述符时F1值为0.30,证实了创始人叙事的独立预测价值。特征分析显示,包括形容词、行话和流行词在内的炒作标记的优化密度与更高的退出概率相关,而过多的陈述或名称长度则会降低该概率。本研究还为风险投资应用引入了可量化的“炒作得分(Hyping Score)”,表明在信息高度不对称的情况下,初创企业的叙事框架能为预测退出提供可衡量的信号。
英文摘要
This study shows that textual descriptors alone can predict early-stage startup success, defined as Exit, without relying on contextual, financial, or human capital variables. Using venture capital-curated datasets covering 7,419 startups over 20 years, the research isolates text-based framing variables and engineers 850 features via startup narrative mapping. Data subsets and vector embeddings are evaluated for statistical significance, followed by supervised machine learning experiments across six models. Binary Exit prediction using Logistic Regression attains an F1 of 0.48 with 0.55 recall using all features (excluding embeddings), and an F1 of 0.26 with 0.59 recall using textual descriptors only (including embeddings). Feature analysis indicates that optimized densities of hyping markers such as adjectives, jargon, and buzzwords are associated with higher Exit probability, while excessive statement or name length is associated with lower probability. The study also introduces a quantifiable Hyping Score for potential application in venture screening. Findings indicate that startup framing can serve as standalone predictor of economic outcomes, in high-information-asymmetry investment environments.