arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20384cs.AI

基于线性判别树集成的可解释多模态分类

Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles

Mojtaba Moattari

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对多模态分类器需平衡准确率与可解释性的需求,提出基于线性判别树集成的框架,在多模态情感行为分类任务中实现了优于基线模型的性能与更高的人类标注一致性。

中文摘要 AI 辅助

融合文本、音频、视觉等异质流的多模态情感与行为分类器,需同时达到有竞争力的准确率,并生成人类可理解的决策线索解释——这一双重目标是当前高容量模型(尤其是Transformer)仅部分解决的问题。Transformer虽能实现较强的预测性能,但其分布式表示与深度非线性特性,使得难以给单个多模态特征分配有意义的重要性权重,限制了其在临床情感监测、教育评估等对信任敏感的应用中的使用。本研究通过开发基于树集成的框架解决该缺口,平衡准确率与可解释性。该框架将每个模态编码为token,提取并聚类概念以降维,将融合后的模态输入树集成分类器,并通过新型改进特征重要性指标解释趋势。该改进指标可降低二元分类任务中负类的影响,从而提升指标或标记检测效果。所提出的树集成模型包括线性判别树(LDT)、线性判别森林(LDF)和线性判别AdaBoost(LDAB),相较于多模态Transformer,其F1-mod提升4.3%;相较于主要的可解释多模态基线模型——可解释多模态路由(IMR),其准确率提升3.0%。所提出的多模态特征重要性提取显著跨模态概念,其人类标注者一致性得分远高于默认特征重要性:在IEMOCAP数据集上为62.2% vs. 43.2%,在CMU-MOSI数据集上为46.7% vs. 32.1%。

英文摘要

Multimodal affect and behaviour classifiers that fuse heterogeneous text, audio, and visual streams must simultaneously achieve competitive accuracy and produce human-understandable explanations of the cues driving their decisions -- a dual objective that current high-capacity models, notably Transformers, only partially address. While Transformers attain strong predictive performance, their distributed representations and deep nonlinearity make it difficult to assign meaningful importance weights to individual multimodal features, limiting their use in trust-sensitive applications such as clinical affect monitoring and educational assessment. We address this gap by developing a framework based on tree-based ensembles that balances accuracy and interpretability. The framework encodes each modality into tokens, extracts and clusters concepts to reduce dimensionality, routes the fused modalities through tree-based ensemble classifiers, and interprets trends using a novel modified feature importance metric. The modified importance reduces the influence of the negative class in binary classification tasks, thereby improving indicator or marker detection. The proposed tree-based ensembles -- Linear Discriminant Tree (LDT), Linear Discriminant Forest (LDF), and Linear Discriminant AdaBoost (LDAB) -- achieve F1-mod gains of 4.3\% over the Multimodal Transformer and accuracy gains of 3.0\% over the primary interpretable multimodal baseline, Interpretable Multimodal Routing (IMR). The proposed multimodal feature importance extracts salient inter-modal concepts with substantially higher human-annotator agreement scores than default feature importance (62.2\% vs.\ 43.2\% on IEMOCAP; 46.7\% vs.\ 32.1\% on CMU-MOSI).

↑