Google Play 用户评论情感指数的统计基础:信号融合、收缩、分布验证与动态平滑
Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing
浏览论文内容
中文总结 AI 辅助
本文为 Google Play 用户评论构建统计情感指数,融合星级与文本情感,经收缩、分布验证及卡尔曼平滑,证明其数学基础,并示例说明文本投诉可降低纯星级评分。
中文摘要 AI 辅助
我们为 Google Play 用户评论开发了一个统计上明确的情感指数,并建立了支持其构建的数学结果。归一化的星级评分和文本情感分数被视为潜在评论效价的噪声度量,并通过协方差感知的逆方差加权进行融合。评论级估计使用有界的帮助性和时效性权重进行聚合,然后使用估计的精度而非任意的评论数量阈值向总体均值收缩。应用级评分直方图为在不同 API 排序顺序下返回的样本提供了分布诊断;由于星级评分是离散的,因此不使用经典连续 Kolmogorov-Smirnov 临界值。局部水平状态空间模型和卡尔曼滤波器提供了去噪的时间趋势。完整证明涵盖了 BLUE 和高斯最大似然结果、高斯共轭收缩、Glivenko-Cantelli 和 Donsker 定理、通过 delta 方法的计数变换以及精确高斯卡尔曼滤波。一个包含三条评论的示例展示了文本投诉如何能实质性地降低看似完美的纯星级评分。
英文摘要
The decision problem is whether valence among the users who write reviews of an app has changed enough to warrant investigation of a release, outage, or support issue. We propose a statistical specification for that decision, applied to written Google Play reviews collected under a declared locale, language, sort order, retrieval cap, and window rule. That endpoint returns a selected collection rather than a random sample, so the index describes review writers under the declared protocol and not all app users. What is new is not a sentiment classifier but an auditable interface: an ordinal threshold model is the preferred representation of star ratings, a calibrated two-signal generalized least-squares estimator is retained as an interpretable linear approximation when held-out labels support its conditional-unbiasedness assumption, the estimand is declared before the estimator, and every downstream module--uncertainty, diagnostics, and regularization--is tied to that declaration. The framework separates the latent average of a fixed collected set from the mean of a protocol-defined population of review writers, so measurement and between-review variation are not conflated, and treats helpfulness and recency weights as a different policy target. Shrinkage, cross-app empirical Bayes, histogram and dependence diagnostics, and dynamic smoothing are specified as separate modules with defaults and mandatory sensitivity analyses. A reference implementation, a declared 6000-review three-app collection, and a synthetic stress test exercise the displayed equations: fusion reduced error in clean and small-window scenarios, coordinated contamination caused severe bias and interval failure, and the collected sample reproduces the abstention rules. No real-review validation is claimed.
发表机构
- Bulacan State University(布拉坎州立大学)
机构由 AI 辅助整理,请以论文原文为准。