发表机构
Instituto de Informática, Universidade Federal de Goiás; Instituto de Ciências Matemáticas e de Computação, Universidade de São Paulo; Escola Superior de Agricultura Luiz de Queiroz da Universidade de São Paulo; Faculdade de Saúde Pública, USP(戈亚斯联邦大学信息学院; 圣保罗大学数学与计算机科学学院; 圣保罗大学路易斯·德·奎罗斯高等农业学院; 圣保罗大学公共卫生学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以圣保罗市76个区为样本,通过整合多源数据,利用8种机器学习分类器分析发现,食物环境指标可区分区级社会脆弱性,其中健康与不健康食品机构密度是关键特征,XGBoost表现最优,但研究存在泛化局限。
AI 中文摘要
城市食物环境可能反映更广泛的社会经济不平等,但巴西城市的区级证据仍然有限。本研究检验圣保罗市96个区中,食品零售和街头市场可及性指标是否能区分不同水平的社会脆弱性。我们开展了探索性横断面生态分析,整合了圣保罗社会脆弱性指数(IPVS)、来自Relação Anual de Informações Sociais(RAIS)的机构记录,以及来自CAISAN的街头市场数据。将普查区信息汇总至区级水平,排除20个无IPVS分类的区,最终得到76个观测值。结局变量区分IPVS 1级的区与2至7级的区,预测变量包括健康与不健康食品机构的密度、街头市场数量,以及销售新鲜或天然食品的机构的可及性。使用留一法交叉验证评估8种传统机器学习分类器,报告的平均F分数范围为0.62至0.75,其中XGBoost取得最高值。在随机森林模型中,健康与不健康食品机构的密度共同占基于杂质的总特征重要性的约60%。这些发现表明,公开可用的食物环境指标包含与区级社会脆弱性分布相关的信息,但较小的生态样本、类别不平衡、结局二值化及横断面设计限制了预测泛化性,无法进行因果或家庭层面的解释。
英文摘要
Urban food environments may reflect broader socioeconomic inequalities, but district-level evidence remains limited in Brazilian cities. This study examined whether indicators of food retail and street-market availability discriminate between levels of social vulnerability across the 96 districts of São Paulo. We conducted an exploratory cross-sectional ecological analysis integrating the São Paulo Social Vulnerability Index (IPVS), establishment records from the Relação Anual de Informações Sociais (RAIS), and street-market data from CAISAN. Census-sector information was aggregated at the district level. Twenty districts without an IPVS classification were excluded, resulting in 76 observations. The outcome distinguished districts classified as IPVS level 1 from those classified as levels 2--7. Predictors described the densities of healthy and unhealthy food establishments, the number of street markets, and the availability of establishments selling fresh or in natura food. Eight conventional machine-learning classifiers were evaluated using leave-one-out cross-validation. Reported mean F-scores ranged from 0.62 to 0.75, with XGBoost obtaining the highest value. In the Random Forest model, the densities of healthy and unhealthy food establishments jointly accounted for approximately 60% of the total impurity-based feature importance. These findings indicate that publicly available food-environment indicators contain information associated with the district-level distribution of social vulnerability. However, the small ecological sample, class imbalance, outcome binarization, and cross-sectional design limit predictive generalization and preclude causal or household-level interpretations.