发表机构
TH Köln; Leiden University; Toyota Racing(科隆应用科学大学; 莱顿大学; 丰田赛车)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对深度不平衡回归评估中的三个盲点,通过多模态虚拟传感基准、引入bMASE指标及多随机种子再评估,揭示尾部区域不稳定等隐藏失效模式,为构建可靠捕获罕见目标工况的回归系统提供可复现基础。
AI 中文摘要
深度不平衡回归(DIR)解决了回归模型的一个常见失效模式:目标分布高度不均匀,导致模型在目标值密集的区域表现最佳,即使在整个目标范围内都需要可靠性能。尽管方法学进展迅速,DIR评估仍受三个盲点制约:它由基于图像的基准主导,其标准的多/中/少样本协议具有诊断性但非决策完备性,且尾部区域在不同随机种子下的稳定性尚未被系统评估。我们沿着这三个轴重新审视DIR评估。首先,我们通过在多模态虚拟传感基准(\textsc{MuViS})上评估DIR来拓宽数据领域,该基准包含跨越六个物理领域的九个时间序列外生回归任务,其中罕见的目标值通常对应操作上有意义的工况。其次,我们采用平衡平均绝对误差(\emph{bMAE})并引入平衡平均绝对缩放误差(\emph{bMASE}),这是一种用于跨方法和数据集进行决策完备比较的尺度归一化指标。第三,通过对六种代表性DIR方法在多个随机种子下的重复再评估,我们表明DIR所针对的尾部区域对种子级变异性表现出特别高的敏感性。我们的结果表明,标准虚拟传感模型表现出被全局MAE掩盖的显著尾部退化,现有DIR方法可以改善平衡性能,但向多模态时间序列数据的迁移不均匀,且尾部区域不稳定性在当前DIR评估实践下仍是一个很大程度上隐藏的失效模式。总之,这些发现和我们的公开代码为未来DIR研究提供了可复现的基础,以构建能像捕获常见目标工况一样可靠地捕获罕见目标工况的回归系统。
英文摘要
Deep Imbalanced Regression (DIR) addresses a common failure mode of regression models: target distributions are highly non-uniform, causing models to perform best in densely populated target regions even when reliable performance is required across the full target range. Despite rapid methodological progress, DIR evaluation remains constrained by three blind spots: it is dominated by image-based benchmarks, its standard many-/medium-/few-shot protocol is diagnostic but not decision-complete, and tail-region stability across random seeds has not been systematically evaluated. We revisit DIR evaluation along these three axes. First, we broaden the data domain by evaluating DIR on a multimodal virtual sensing benchmark (\textsc{MuViS}) with nine time-series extrinsic regression tasks across six physical domains, where rare target values often correspond to operationally meaningful regimes. Second, we adopt balanced MAE (\emph{bMAE}) and introduce balanced Mean Absolute Scaled Error (\emph{bMASE}), a scale-normalized metric for decision-complete comparison across methods and datasets. Third, through a repeated reevaluation of six representative DIR methods across multiple random seeds, we show that the tail regions targeted by DIR exhibit particularly high sensitivity to seed-level variability. Our results show that standard virtual-sensing models exhibit substantial tail degradation hidden by global MAE, that existing DIR methods can improve balanced performance but transfer unevenly to multimodal time-series data, and that tail-region instability remains a largely hidden failure mode under current DIR evaluation practice. Together, these findings and our publicly available code provide a reproducible basis for future DIR research toward regression systems that capture rare target regimes as reliably as common ones.