arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BiasFlow: 用于虚假特征依赖的几何监测与骨干正则化

BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance

Haojin Deng, Zhiping Lin, Yimin Yang

arXiv 2610.06846首次发表:更新:

发表机构

Western University; Nanyang Technological University; Vector Institute(西安大略大学; 南洋理工大学; 向量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BiasFlow通过监测质心几何并施加BFR正则化,减少虚假特征依赖,提升最差组准确率,在多个基准上验证了有效性。

AI 中文摘要

最差组准确率(WGA)评估训练后的预测器,但并未刻画其冻结的骨干网络在学习新头部时的行为。我们引入了BiasFlow,一个基于钩子的工具包,用于监测类属性质心对齐(IBMI)、类内质心分离(W-IBMI)以及特征投影敏感性。IBMI受类属性相关性的混淆,并非因果特征依赖的度量。我们将这些诊断与BiasFlow正则化(BFR)配对,BFR是一种有监督的、可组合的类条件质心对齐惩罚。W-IBMI验证了BFR所优化的量;它依赖于尺度,且不能独立确立属性移除。在报告的小规模基准上,添加BFR可改善或保持平均WGA,在UrbanCars上增益高达+26.0个百分点。主要的独立压力测试冻结CelebA-Std骨干,并在有偏数据上训练新头部:BFR+GroupDRO将WGA从40.7%提升至64.1%,而Male探针准确率从92.5%下降至72.2%。属性信息仍然可恢复,跨任务结果好坏参半。一项受控的合成水印ImageNet实验在匹配训练下额外将水印偏移准确率提高了+23.0个百分点。这些结果支持在测试协议内,将质心几何和对有偏头部重训的抵抗性评估与WGA并列进行。

英文摘要

Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen backbone behaves when a new head is learned. We introduce BiasFlow, a hook-based toolkit for monitoring class-attribute centroid alignment (IBMI), within-class centroid separation (W-IBMI), and feature-projection sensitivity. IBMI is confounded by class-attribute correlation and is not a measure of causal feature reliance. We pair these diagnostics with BiasFlow Regularization (BFR), a supervised, composable class-conditional centroid-alignment penalty. W-IBMI verifies the quantity BFR optimizes; it is scale dependent and does not independently establish attribute removal. Across the reported small-scale benchmarks, adding BFR improves or preserves mean WGA, with gains up to +26.0 pp on UrbanCars. The principal independent stress test freezes CelebA-Std backbones and trains fresh heads on biased data: BFR+GroupDRO improves WGA from 40.7% to 64.1%, while Male probe accuracy decreases from 92.5% to 72.2%. Attribute information remains recoverable, and cross-task results are mixed. A controlled synthetic-watermark ImageNet experiment additionally improves watermark-shift accuracy by +23.0 pp under matched training. These results support evaluating centroid geometry and resistance to biased head retraining alongside WGA, within the tested protocols.

Comments19 pages, 7 figures, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑