AI 中文总结
本文提出首个基于梯度提升树的CCA方法TreeCCA,兼具非线性精度与可解释性,在多类基准上性能优于或相当Deep CCA等方法,成本更低且可提供数据解释。
AI 中文摘要
梯度提升树(GBT)在表格机器学习中占据主导地位,但典型相关分析(CCA)一直依赖线性或神经编码器。我们提出TreeCCA,这是首个将梯度提升树集成作为CCA编码器端到端训练的方法,继承了其即插即用的可靠性:无需架构设计、采用熟悉的超参数,默认设置下性能强劲。技术实现是Eckart-Young(EY)损失,它提供了逐样本闭式梯度,可作为自定义目标直接嵌入任何标准GBT库(XGBoost、LightGBM)。TreeCCA是首个兼具非线性精度与原生可解释性的CCA方法:每个树分裂选择一个特征,增益重要性可揭示哪些输入驱动跨视图相关性,且无需额外成本。我们在合成基准上验证了这些特性,TreeCCA在Signed Power数据集上为2.61,优于Deep CCA的2.43;在Hermite数据集上为2.93,优于Deep CCA的2.89。在零线性跨视图协方差的稀疏基准中,TreeCCA在p=50时Precision@S=1.00,恢复了真实支持,而PMD未找到任何信号。在UCI HAR传感器融合基准上,TreeCCA精度与Deep CCA相当,但成本降低5倍,且XGBoost增益重要性直接验证了关于数据的物理动机假设——神经编码器无法轻易提供此类解释。在五个流行的表格多视图数据集上,TreeMCCA在非线性相关性提取和下游分类精度上始终与线性CCA相当或更优。
英文摘要
Gradient-boosted trees dominate tabular machine learning, yet canonical correlation analysis has always relied on linear or neural encoders. We propose \textbf{TreeCCA}, the first method to train gradient-boosted tree ensembles end-to-end as CCA encoders, inheriting their plug-and-play reliability: no architecture design, familiar hyperparameters, and strong performance with defaults. The technical enabler is the Eckart-Young (EY) loss, which supplies closed-form per-sample gradients that slot directly into any standard GBT library (XGBoost, LightGBM) as a custom objective. TreeCCA is the first CCA method to combine nonlinear accuracy with native interpretability: every tree split selects one feature, so gain importances reveal which inputs drive cross-view correlation at no extra cost. We demonstrate these properties on synthetic benchmarks, where TreeCCA matches or exceeds Deep CCA (2.61 vs.\ 2.43 on Signed Power; 2.93 vs.\ 2.89 on Hermite), and on a sparse benchmark with zero linear cross-view covariance, where TreeCCA recovers the true support with $\text{Precision@}S = 1.00$ at $p=50$ while PMD finds no signal. On the UCI HAR sensor-fusion benchmark, TreeCCA achieves comparable accuracy to Deep CCA at $5\times$ lower cost, while XGBoost gain importances directly validate a physics-motivated hypothesis about the data --- an interpretation not readily available with neural encoders. Across five popular tabular multi-view datasets, TreeMCCA consistently matches or exceeds linear CCA in both nonlinear correlation extraction and downstream classification accuracy.