arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分布式部署下医学模型性能的可靠性测试

Reliability Testing of Medical Model Performance under Distributed Deployment

Yifei Wang, Xiaohan Zhang, Youtao Ding, Tianlin Li, Xiaoyu Zhang, Yida Yang, Li Pan

arXiv 2609.36525首次发表:更新:

发表机构

Shanghai Jiao Tong University; Beihang University; Nanyang Technological University; Tongji University(上海交通大学; 北京航空航天大学; 南洋理工大学; 同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对分布式部署导致医学模型输出与集中式评估不一致的问题,提出测试框架与分布式执行敏感基准,实验显示执行变化可致输出分歧,旨在将评估扩展到部署一致性。

AI 中文摘要

分布式推理已成为在实际延迟、内存和吞吐量约束下部署医学模型不可或缺的一部分。尽管现代框架通过张量并行、混合精度、内核融合和多设备通信提高了服务效率,但它们通常被假定能保持集中式HuggingFace评估期间观察到的行为。这一假设造成了评估与部署之间的不匹配:模型可能通过离线评估,但在执行栈改变后产生不同的输出。为解决这一不匹配,我们提出了一种测试框架和一个改进的、对分布式执行敏感的医学模型基准,该基准在集中式HuggingFace参考和匹配的分布式部署下评估相同的检查点和输入。跨语言、视觉和多模态医学模型的广泛实验表明,执行变化可产生可测量的输出分歧。在支持的视觉设置中,单模态模型的测试成功率范围为0.21至0.43,多模态模型为0.32至0.98。该基准旨在将医学模型评估从能力和安全性扩展到评估-部署一致性。

英文摘要

Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑