发表机构
University of the Aegean; Athens University of Economics and Business; University College London; London School of Economics(爱琴海大学; 雅典经济与商业大学; 伦敦大学学院; 伦敦政治经济学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出贝叶斯正例-未标注分类框架,结合高斯过程与多层网络检测税务欺诈,在希腊能源市场案例中验证了高排序性能并识别高风险企业。
AI 中文摘要
税务欺诈仍然是全球公共税务机关面临的核心挑战,仅在欧盟每年造成的财政损失估计高达1万亿欧元。欺诈性企业故意操纵报告数据以掩盖其活动。商业网络可以提供补充信息,并揭示仅凭协变量可能遗漏的欺诈行为。我们提出了一种贝叶斯正例-未标注(PU)分类框架,将企业层面的协变量与多层网络信息相结合。作为一项政策相关的案例研究,我们将该框架应用于希腊能源市场受监管细分领域的企业,该领域面临消费税逃税风险。已确认的欺诈标签仅存在于已审计案例中,其余观测值未被标注,而非已验证合规。我们开发了一种贝叶斯高斯过程(GP)分类器,通过专家乘积(PoE)构造将协变量与多层网络信息整合,并纳入漏检概率以处理遗漏的正例。我们在锚定条件下推导了漏检率的识别界限,并表明对于固定的潜在风险函数,共同的漏检率影响校准但不影响排序。模拟实验表明,相对于PU和基于网络的替代方法,该方法具有强大的排序性能,同时说明了估计漏检率的难度。在实际审计数据中,该方法识别出具有后验不确定性的高风险企业,估计未检测到的欺诈案件数量,并确定风险是由协变量、网络还是两者共同驱动。在排名最高的六家企业中,有两家已被独立重新检查并均确认为不合规,为这些案例提供了独立验证。其余四家已被税务机关选中进行后续检查。
英文摘要
Tax fraud remains a central challenge for public revenue authorities worldwide, imposing fiscal losses estimated to reach up to 1 trillion euros annually in the EU alone. Fraudulent firms intentionally manipulate reported figures to conceal their activity. Business networks can provide complementary information and reveal fraud that covariates alone may miss. We propose a Bayesian positive-unlabeled (PU) classification framework that combines firm-level covariates with multilayer network information. As a policy relevant case study, we apply the framework to firms in a regulated segment of the Greek energy market, a sector exposed to excise tax evasion. Confirmed fraud labels exist only for audited cases, while the remaining observations are unlabeled rather than verified compliant. We develop a Bayesian Gaussian process (GP) classifier integrating covariates with multilayer network information through a Product of Experts (PoE) construction and incorporating a nondetection probability for missed positives. We derive identification bounds for the nondetection rate under an anchor condition and show that, for a fixed latent risk function, a common nondetection rate affects calibration but not ranking. Simulations show strong ranking performance relative to PU and network-based alternatives, while illustrating the difficulty of estimating the nondetection rate. In real audit data, the method identifies high risk firms with posterior uncertainty, estimates the number of undetected fraudulent cases, and identifies whether risk is driven by covariates, networks, or both. Of the six highest ranked firms, the two that had already been reinspected independently were both confirmed as noncompliant, providing an independent validation of these cases. The remaining four were selected by the tax authority for follow-up inspection.