arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40019nucl-thhep-ph

泊松分布数据及比值的卡方变体比较

Comparison of chisquare variants for Poisson-distributed data and ratios

Mate Csanad, Yan Huang, Daniel Kincses, Marton I. Nagy, Barnabas Porfy

首次发表
浏览论文内容

中文总结 AI 辅助

本文比较多种卡方变体对泊松分布数据及比值拟合的偏差,发现Neyman型有约3/λ的相对偏差,而似然法无偏,推荐用于飞米尺度关联函数分析。

中文摘要 AI 辅助

在高能物理和重离子物理中,对分箱的泊松分布数据进行拟合十分常见,而在关联函数测量中尤为精细,因为此时拟合的观测量是两个直方图的比值。我们在大型玩具数据集上比较了多种拟合优度估计量:Neyman和Pearson的χ²、它们的Yates连续性校正版本、方差平移1/2的Neyman χ²、泊松对数似然,以及将信号和参考直方图均视为泊松分布的关联函数似然。对于平均占据数为λ的单个直方图,Neyman χ²将箱内容低估约一个计数,而Pearson χ²则高估约半个计数;将方差平移1/2仅在1/λ阶上改变Neyman结果,且只有对数似然能无偏地恢复均值。对于两个直方图的比值,Neyman型估计量在相对意义上约有-3/λ的偏差(在λ=100时为3%),而Pearson和基于似然的拟合在所研究的对称配置中是无偏的。Yates校正基本不改变拟合值,但会系统性地降低χ²,导致不切实际的高置信水平。我们推导出能重现所有观测偏差的简单解析表达式。由于这些偏差不随箱数减少,而统计不确定性却会减少,因此它们可能主导高统计量、精细分箱测量的不确定性。因此,我们推荐基于似然的拟合,特别是关联函数似然,用于费米子源成像及类似的比值分析。

英文摘要

Fits to binned, Poisson-distributed data are common in high-energy and heavy-ion physics, and they are particularly delicate in correlation function measurements, where the fitted observable is the ratio of two histograms. We compare several goodness-of-fit estimators on large toy data sets: the Neyman and Pearson $χ^2$, their Yates continuity-corrected versions, a Neyman $χ^2$ with the variance shifted by $1/2$, the Poisson log-likelihood, and the correlation function likelihood in which both the signal and the reference histogram are treated as Poisson distributed. For a single histogram with mean occupancy $λ$, the Neyman $χ^2$ underestimates the bin content by approximately one count, while the Pearson $χ^2$ overestimates it by approximately half a count; shifting the variance by $1/2$ changes the Neyman result only at order $1/λ$, and only the log-likelihood recovers the mean without bias. For the ratio of two histograms, Neyman-type estimators are biased by approximately $-3/λ$ in relative terms (3\% at $λ=100$), whereas the Pearson and likelihood-based fits are unbiased in the symmetric configuration studied here. The Yates correction leaves the fitted values essentially unchanged but systematically deflates the $χ^2$, resulting in unrealistically high confidence levels. Simple analytic expressions are derived that reproduce all observed biases. Since these biases do not decrease with the number of bins, whereas the statistical uncertainties do, they can dominate the uncertainty of high-statistics, finely binned measurements. We therefore recommend likelihood-based fits, in particular the correlation function likelihood, for femtoscopic and similar ratio analyses.

发表机构

  • ELTE Eötvös Loránd University(罗兰大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑