验证什么有助于客户退货时机:一种针对条件信号的筛选与确认测试,以及为何衰减几乎足够
Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough
浏览论文内容
中文总结 AI 辅助
本研究提出筛选与确认协议及无模型上限工具,验证了连续时间衰减几乎可满足客户退货时机预测,而额外添加的条件信号多为冗余或有害。
中文摘要 AI 辅助
从业者不断为客户退货模型添加各类信号(终身价值、品类、近期性/频率、日历、地理位置),而时间点过程(TPP)文献也随之跟进,提出了协变量及外部协变量条件强度模型。但这些信号是否真能改善时机预测?又该如何验证?仅当模型本可检测到某信号时,“特征X无帮助”这一零假设才有意义。我们为此做出两项贡献——一套方法与一种衡量标准,以可靠回答上述问题:(i)筛选与确认协议,用于验证候选信号是否提升TPP的事件时机似然:阳性对照植入已知强度的耦合关系,并确认模型可将其恢复,因此真实数据中的零假设可解读为“无信号”而非“方法效力弱”;该对照已针对类别型与连续型编码,以及真实时钟驱动数据集(纽约出租车小时时段)完成验证。(ii)无模型上限,量化客户退货时机中可被点预测的最小占比(来自任何协变量的差距方差仅为个位数百分比;退货几乎无记忆性)。借助上述工具,我们在三个公开基准数据集(Amazon、Taobao、RetailRocket)及一个真实市场平台(Thumbtack)上验证了清晰结果:事件间时钟——连续时间衰减(长期以来优于冻结强度模型)几乎足够,而该领域不断添加的条件信号在此基础上是冗余或有害的(公开基准上统计零效应,负对数似然(NLL)至多为0.06;真实平台上零效应至轻度有害)。我们并非声称发现衰减有帮助,而是贡献了将“条件信号无帮助”转化为可验证、已认证表述的工具——以及对我们遇到并修正的读出/泄漏陷阱的诚实评估说明。
英文摘要
Practitioners enrich customer-return models with ever more signals (lifetime value, category, recency/frequency, calendar, geography), and the temporal-point-process (TPP) literature follows suit with covariate- and external-covariate-conditioned intensities. But does any of it improve the timing, and how would you know? A null ("feature X doesn't help") is only meaningful if the model could have found a signal. We make two contributions--a method and a measurement--to answer this credibly. (i) A screen-and-confirm protocol that certifies whether a candidate signal improves a TPP's event-timing likelihood: a positive control plants a coupling of known strength and confirms the model recovers it, so a real-data null can be read as "no signal" rather than "weak method." The control is validated for categorical and continuous encodings, and on a real clock-driven dataset (NYC taxi hour-of-day). (ii) A model-free ceiling quantifying how little of customer-return timing is point-predictable at all (a single-digit percentage of gap variance from any covariate; returns are near-memoryless). With these we certify a clean result on three public benchmarks (Amazon, Taobao, RetailRocket) and a real marketplace (Thumbtack): the inter-event clock--continuous-time decay, long known to beat frozen-intensity models--is nearly sufficient, and the conditioning the field keeps adding is redundant or harmful on top of it (statistically null on the public benchmarks, at most 0.06 NLL; null to mildly harmful on the marketplace). We do not claim to discover that decay helps; our contribution is the tools that turn "conditioning doesn't help" into a checkable, certified statement--plus an honest-evaluation account of the read-out/leakage pitfalls we hit and retracted.
发表机构
- Thumbtack, Inc.(Thumbtack公司)
机构由 AI 辅助整理,请以论文原文为准。