AI 中文总结
研究非契约性商业中企业统计客户数量的问题,指出“购买直到死亡”模型混淆不同量,通过航位推算得隐含计数,以七年面板数据为例说明其不确定性,提出报告可审计期限计数等补救措施。
AI 中文摘要
在非契约性商业中,企业面临着了解实际客户数量的挑战,因为客户可能在未表明离开的情况下停止购买。“购买直到死亡”模型通过估计每个客户存活的概率(称为P(存活))来解决这个问题,该概率用于每个主要的软件工具中。我们表明这种做法混淆了两个不同的量。在贝塔 - 几何族中,P(存活)是有限期重复购买概率的可观察族的无限期极限。每个有限期估计,如12个月内重复购买的概率,是对可验证事件的预测。无限期极限只能通过外推得到,即通过航位推算获得的客户数量。因此,隐含的计数只是部分确定的:已实现的回头客是下限,估计惯例决定了高于该下限的报告点估计。在一个包含31,683个客户的七年面板上,具有几乎相同可观察预测的规格估计存活客户数量在3,654到27,734之间,相差7.6倍;仅一个默认软件加权参数就使计数波动42%;五年后的购买从下方证伪了最大似然计数。这些模式在CDNOW基准上也有体现,相差2.4倍。实践中所谓的错误校准大多是类别错误:P(存活)总和比已实现的18个月回头客多2.25倍,而同一模型自己的18个月预测误差仅为1.18倍。补救措施是报告可审计的期限计数,在规定期限内估计回头概率,在不同评分日期和期限进行审计,随着队列变化重新校准,如果要报告总数,应作为区间而非点。
英文摘要
Firms in non-contractual commerce face the challenge of knowing how many customers they actually have because customers can stop buying without ever saying they have left. Buy-Till-You-Die models address this by estimating each customer's probability of being alive, a quantity called P(alive) and used in every major software tool for dashboards, churn, customer equity, and enterprise valuation. We show this practice confounds two distinct quantities. Within the beta-geometric family, P(alive) is the infinite-horizon limit of an observable family of finite-horizon repeat-purchase probabilities. Every finite-horizon estimate, such as the probability of repeat purchase within 12 months, is a forecast of a verifiable event. The infinite-time limit can only be reached by extrapolation, a customer count obtained by dead reckoning. The implied count is therefore only partially identified: realized returners are the lower bound, and estimation conventions determine the reported point estimate above that. On a seven-year panel of 31,683 customers, specifications with nearly identical observable forecasts estimate the number of alive customers anywhere from 3,654 to 27,734, a factor of 7.6; a default software weighting parameter alone swings the count 42 percent; and five years of later purchases falsify the maximum-likelihood count from below. The patterns replicate on the CDNOW benchmark, with a 2.4x spread. Most of what practice calls miscalibration is instead a category error: summed P(alive) overshoots realized eighteen-month returners by 2.25x, while the same model's own eighteen-month forecast errs by just 1.18x. The remedy is to report an auditable horizon count, estimate return probabilities at a stated horizon, audit them across scoring dates and horizons, recalibrate as cohorts drift, and report the total count, if at all, as an interval rather than a point.