发表机构
Fraunhofer IGD; Technische Universität Darmstadt(弗劳恩霍夫计算机图形学研究所; 达姆施塔特工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对网络流量分类模型仅追求预测性能而忽视语义可信度的问题,提出一个结合数据、模型、可解释性、可视化与专家推理的以人为中心框架,以支持语义验证,提升模型的稳健性与可信度。
AI 中文摘要
机器学习(ML)已成为网络流量分类的主导方法,实现了非常高的预测性能。然而,只有当模型学习到语义上有意义且可信的模式,而非利用虚假相关性时,模型才具有价值。传统的评估实践主要评估预测性能,因此模型是否依赖语义上有意义的模式仍然未知。为应对这些挑战,我们将知识生成框架适配到网络流量分类中。该适配框架结合了数据、机器学习模型、可解释性、可视化和专家推理,以支持对模型行为和数据预处理的迭代探索、验证与改进。该框架基于文献发现、基准数据集分析、基于XAI的流量分类实践经验以及专家反馈,为语义模型验证提供实用指导。通过将预测性能与语义验证和人类专业知识相结合,所提出的框架支持开发不仅准确而且稳健可信的网络流量分类模型。
英文摘要
Machine learning (ML) has become the dominant approach for network traffic classification, achieving very high predictive performance. However, a model is only valuable if it learns semantically meaningful and trustworthy patterns rather than exploiting spurious correlations. Conventional evaluation practices predominantly assess predictive performance. Consequently, whether the model relies on semantically meaningful patterns remains unknown. To address these challenges, we adapt the knowledge generation framework for network traffic classification. The adapted framework combines data, ML models, explainability, visualization, and expert reasoning to support the iterative exploration, verification, and refinement of model behavior and data preprocessing. The framework is grounded in findings from the literature, benchmark dataset analyses, practical experience with XAI-based traffic classification, and expert feedback, providing practical guidance for semantic model validation. By complementing predictive performance with semantic validation and human expertise, the proposed framework supports the development of network traffic classification models that are not only accurate but also robust and trustworthy.