发表机构
School of Digital and Physical Sciences, University of Hull(赫尔大学数字与物理科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种整合LLM引导设计与帕累托优化的以人为中心XAI选择框架,用于TinyML边缘设备部署,经皮肤病变分类任务验证可识别帕累托有效权衡。
AI 中文摘要
边缘人工智能(Edge AI)可将AI模型直接部署在本地边缘设备上,此类部署受严格资源约束,尤其在需本地及时推理的临床应用中。在此场景下,可解释人工智能(XAI)可作为人机交互界面,助力医护人员与患者理解模型预测并做出知情决策。为实现该作用,TinyML部署的XAI方法选择可被构建为以人为中心的多目标设计问题,需同时考量利益相关者的定性偏好、解释质量及基于代理的部署成本。我们提出的框架整合了大语言模型(LLM)引导的设计界面,该界面可将利益相关者的定性偏好映射为候选XAI方法,随后进行确定性可行性过滤与基于帕累托的优化。该框架揭示了解释保真度、稳定性及基于代理的部署成本间的权衡关系,同时明确其对解释质量与预估部署可行性的影响。针对皮肤病变分类任务的概念验证评估表明,该框架可系统比较候选XAI方法并识别帕累托有效权衡。本次评估涵盖计算选择阶段,而物理MCU部署与实证人类专家验证不在本研究范围内。
英文摘要
Edge Artificial Intelligence (Edge AI) enables the deployment of AI models directly on local edge devices, while such deployments are subject to strict resource constraints, particularly in clinical applications requiring local and timely inference. In such contexts, explainable artificial intelligence (XAI) can serve as a human-AI interface intended to support healthcare professionals' and patients' understanding of model predictions and informed decision-making. To fulfill this role, XAI method selection for TinyML deployments can be formulated as a human-centered multi-objective design problem that jointly considers qualitative stakeholder preferences, explanation quality, and proxy-based deployment cost. We propose a framework that integrates a large language model (LLM)-guided design interface that maps qualitative stakeholder preferences to candidate XAI methods, followed by deterministic feasibility filtering and Pareto-based optimization. The framework exposes trade-offs among explanation fidelity, stability, and proxy-based deployment cost while characterizing their implications for explanation quality and estimated deployment feasibility. A proof-of-concept evaluation on a skin lesion classification task illustrates how the framework systematically compares candidate XAI methods and identifies Pareto-efficient trade-offs. The present evaluation covers the computational selection stages, while physical MCU deployment and empirical human-expert validation remain outside the scope of this study.