量化治疗型大语言模型(LLMs)的临床安全性与环境影响之间的关系
Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs
浏览论文内容
中文总结 AI 辅助
该研究量化治疗型LLMs的临床安全性与环境影响的关系,发现安全分布高端存在非线性权衡,额外测试时计算未必提升安全,提出动态模型选择等可持续部署方法。
中文摘要 AI 辅助
大型语言模型(LLMs)在心理健康领域的应用引发了关于临床安全性与环境成本之间关系的疑问。本文通过将K-Bench临床安全评分与EcoLogits生命周期评估估算值相结合,对47种受支持的模型配置展开研究,从能源消耗、碳排放、水资源消耗及非生物资源消耗四个维度评估模型性能与环境影响。结果显示,在安全分布的高端存在非线性权衡:临床安全评分每提升2.61个百分点,每百万输出令牌的估算能源消耗约增加60倍。逐行分析进一步表明,额外的测试时计算并未持续提升临床安全性,部分配置下还与更低的临床安全评分相关。这些发现提示,仅依赖更大的模型或额外推理时计算,可能并非提升治疗型AI系统安全性的高效策略。本文探讨了可持续部署的意义,并强调动态模型选择(包括模型级联)是在高风险案例中保持临床性能的同时降低环境影响的潜在方法。
英文摘要
The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine this relationship by combining K-Bench clinical safety scores with EcoLogits life-cycle assessment estimates across 47 supported model configurations. We evaluate model performance and environmental impact across four dimensions: energy use, carbon emissions, water consumption, and abiotic depletion. The results indicate a non-linear trade-off at the upper end of the safety distribution: a 2.61 percentage-point increase in clinical safety score corresponded to an approximately 60-fold increase in estimated energy use per million output tokens. Row-level analyses further suggest that additional test-time compute did not consistently improve clinical safety and, in some configurations, was associated with lower clinical safety scores. These findings suggest that relying solely on larger models or additional inference-time computation may be an inefficient strategy for improving safety in therapeutic AI systems. We discuss the implications for sustainable deployment and highlight dynamic model selection, including model cascading, as a potential approach for reducing environmental impact while preserving clinical performance in higher-risk cases.
发表机构
- University of Isfahan(伊斯法罕大学)
- University of Roehampton(罗汉普顿大学)
- Kivira Health(基维拉医疗)
机构由 AI 辅助整理,请以论文原文为准。