arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11431cs.AI

LLMs作为符号回归中生理学合理性的事后审计者:一项临床医生评估的案例研究

LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study

  • Universidad Complutense de Madrid(马德里康普顿斯大学)
  • Bioinspired Intelligence Ltd.(仿生智能有限公司)
  • Universidad Autónoma de Baja California(下加利福尼亚自治大学)
  • Equ Healthcare(Equ Healthcare 公司)
  • Hospital Universitario de Toledo(托莱多大学医院)
  • Ciudad Real General University Hospital(雷阿尔城综合大学医院)
  • Castilla-La Mancha Health Research Institute (IDISCAM)(卡斯蒂利亚-拉曼恰健康研究所)
  • Hospital Universitario Central de Asturias(阿斯图里亚斯中央大学医院)
  • University of Oviedo(奥维耶多大学)
  • Inst. Investigación Sanitaria Principado de Asturias (ISPA)(阿斯图里亚斯公国卫生研究所)
  • Inst. Universitario Oncología Principado de Asturias (IUOPA)(阿斯图里亚斯公国大学肿瘤研究所)

机构由 AI 辅助整理,请以论文原文为准。

Jorge López-Varela, J. Ignacio Hidalgo, José-Manuel Muñoz, Omar Costilla-Reyes, Esther Maqueda, Jesus Moreno-Fernandez, Tomás González-Vidal, J. Manuel Velasco, Oscar Garnica

AI总结:

本研究探索利用大型语言模型作为事后审计工具,对符号回归生成的表达式进行可解释性和医学合理性排序,经临床医生评估发现其在专家监督下比较审计有效,但不宜自主验证。

AI中文摘要:

遗传编程及其变体(如语法进化)广泛用于符号回归,以从多变量数据中推导数学表达式。除了预测准确性外,模型因其提供可解释性的潜力而受到重视,能够提供将输入变量与结果联系起来的显式方程。然而,实现可解释性和合理性仍然具有挑战性,因为进化出的模型可能复杂或在科学上不一致。在本研究中,我们探讨大型语言模型是否能够帮助改进由进化计算方法生成的符号回归模型的可解释性。基于我们先前使用基于语法的遗传编程估算体脂百分比的工作,我们研究了将LLMs作为后处理工具,根据可解释性和医学合理性分析和排序进化表达式的可能性。四个符号表达式由三个LLMs在三次重复运行中进行分析,所得解释和排序由三位临床医生组成的专家组进行评估。在三个LLMs中,比较模型排序的输出比孤立的术语级解释获得了更有利的临床医生评估。然而,LLMs也产生了生理学和数学上可疑的解释,表明它们更适合在专家监督下进行比较审计,而非自主验证。

英文摘要:

Genetic Programming and its variants, such as grammatical evolution, are widely used in Symbolic Regression to derive mathematical expressions from multivariate data. In addition to predictive accuracy, models are appreciated for their potential to provide interpretability, offering explicit equations that relate input variables to outcomes. However, achieving interpretability and plausibility remains challenging, as evolved models may be complex or scientifically inconsistent. In this study, we explore whether Large Language Models, can assist in improving the explainability of Symbolic Regression models generated by evolutionary computation methods. Building upon our previous work on estimating body fat percentage using grammar-based Genetic Programming , we investigate the use of LLMs as post-processing tools to analyze and rank evolved expressions according to their interpretability and medical plausibility. Four symbolic expressions are analysed by three LLMs over three repeated runs, and the resulting interpretations and rankings are assessed by a panel of three clinicians. Across the three LLMs, comparative model-ranking outputs received more favorable clinician assessments than isolated term-level interpretations. However, the LLMs also produced physiologically and mathematically questionable explanations, indicating that they are better suited to comparative auditing under expert oversight than to autonomous validation.\blfootnote{The present work is an extended version of a paper submitted into a journal.

↑