AI 中文总结
该研究揭示突发人类动态下的事件时间混淆问题,提出形式化偏差分析、诊断方案及工具burstcheck,指出单平台设计无法分离混淆与真实效应,需比较有无事件的相似时段。
AI 中文摘要
数字行为研究常以用户自主选择的时刻(如打开AI助手、点击推荐或访问产品页面)对齐用户,将之后的更高活动解读为事件效应。我们表明这会产生内生时间零点:事件发生在持续的任务时段中,因此对齐的曲线追踪的是时段延续而非对事件的响应。在同用户跨平台网络日志中,AI、购物、新闻、编码和参考事件前均出现广泛的活动增加,且在时间零点前达到峰值。我们最强的测试使用已知无效时间戳(即无任何实际影响的时间戳):在满足严格事件前活动和洗脱标准的5.8%的AI响应中,这些时间戳显示的事件后搜索活动是同用户安慰剂的3.42倍,而真实事件的该比值为4.32倍。已知无效时间戳重现的超额部分比例从可检测活跃时刻的0.56降至安静时刻的-0.04(此时设计未检测到任何效应)。我们将这种时段选择偏差形式化,证明无额外假设时单平台事件窗口无法将其与真实效应分离,并在零效应模拟中展示用户固定效应和粗活动匹配为何失效:该混淆是用户内部且随时间变化的。我们提供诊断方案、公开数据基准及轻量级审计工具burstcheck。用户定时事件可能存在真实效应,但默认情况下事件后活动量无法识别这些效应;研究应比较存在与不存在该事件的相似时段。
英文摘要
Studies of digital behavior often align users at moments they choose, such as opening an AI assistant, clicking a recommendation, or visiting a product page, and interpret higher activity afterward as an event effect. We show how this creates an endogenous time zero: the event occurs during an ongoing task episode, so the aligned curve can trace episode continuation rather than a response to the event. In same-user, cross-surface web logs, AI, shopping, news, coding, and reference events are all preceded by broad activity increases that peak before time zero. Our strongest test uses known-null timestamps that cause nothing. Among the 5.8% of AI responses meeting strict pre-event activity and washout criteria, these timestamps show 3.42 times the post-event search activity of a within-user placebo, compared with 4.32 times for real events. The fraction of excess reproduced by the known null falls from 0.56 at detectably active moments to -0.04 at quiet moments, where the design detects none. We formalize this episode-selection bias, prove that a single-surface event window cannot separate it from a genuine effect without additional assumptions, and show in zero-effect simulations why user fixed effects and coarse activity matching can fail: the confound is within-user and time-varying. We provide a diagnostic protocol, public-data benchmarks, and burstcheck, a lightweight audit tool. User-timed events may have real effects, but post-event volume does not identify them by default; studies should compare similar episodes with and without the event.
Comments20 pages, 7 figures. Includes companion code, simulations, and public-data benchmarks in the source package. Proprietary panel data are not included under the data-provider agreement