发表机构
Brookhaven National Laboratory(布鲁克海文国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对纵向细胞绘画形态学,提出检索增强解释框架,结合多方面证据使LLM生成可审计假设,引入两个定量审计测试,通过实验验证其有效性,产生了可证伪的生物学假设,揭示低剂量率下的适应性表型。
AI 中文摘要
高内涵形态学分析(细胞绘画)可产生细胞状态的敏感、高维特征,但将纵向形态轨迹转化为可解释的生物学内容仍很困难,尤其是对于低剂量率电离辐射等微弱、慢性扰动。大语言模型(LLMs)可将异质证据整合为生物学叙述,但其科学应用需要定量审计。我们提出了一个用于纵向细胞绘画形态学的评估优先、检索增强解释框架,并将其应用于9周的RPE-1时间进程,涵盖五个剂量率(0.003-6.0mGy/hr)。通过稳定的证据标识符,将每周匹配的处理组-对照组形态差异与检索到的扰动邻居、通路背景和文献证据相结合,使LLM能够生成结构化、证据关联的假设,并在保留出处的同时进行分层总结。我们引入了两个定量审计测试:V1引用有效性,验证提示中是否存在引用的证据标识符;V2基于代理的形态兼容性,评估预测的生物学过程与最显著改变的形态特征之间的一致性。在我们的实验中,V1未检测到无效证据引用,而V2显示出有意义的形态兼容性,其随扰动强度增加,并与独立的形态漂移总结呈正相关。该框架产生了可审计、可证伪的生物学假设,包括低剂量率(0.003-0.3mGy/hr)下涉及代谢重编程和蛋白质稳态应激的适应性表型。当前的局限性包括基于代理的评估和缺乏真实机制标签。
英文摘要
High-content morphological profiling (Cell Painting) yields sensitive, high-dimensional signatures of cellular state, but translating longitudinal morphology trajectories into interpretable biology remains difficult, especially for weak, chronic perturbations such as low-dose-rate ionizing radiation. Large language models (LLMs) can synthesize heterogeneous evidence into biological narratives, yet their scientific use requires quantitative auditing. We present an evaluation-first, retrieval-augmented interpretation framework for longitudinal Cell Painting morphology, applied to a 9-week RPE-1 time course across five dose rates (0.003--6.0 mGy/hr). Week-matched treated-control morphology deltas are combined with retrieved perturbation neighbors, pathway context, and literature evidence through stable evidence identifiers, enabling an LLM to generate structured, evidence-linked hypotheses that are hierarchically summarized while preserving provenance. We introduce two quantitative auditing tests: V1 citation validity, which verifies that cited evidence identifiers exist in the prompt, and V2 proxy-based morphology compatibility, which evaluates consistency between predicted biological processes and the most altered morphology features. In our experiments, V1 detected no invalid evidence references, while V2 showed meaningful morphology compatibility that increased with perturbation strength and was positively associated with an independent morphology drift summary. The framework produces auditable, falsifiable biological hypotheses, including an adaptive phenotype involving metabolic reprogramming and proteostatic stress at lower dose rates (0.003--0.3 mGy/hr). Current limitations include proxy-based evaluation and the lack of ground-truth mechanism labels.
CommentsAccepted for publication at ACM BCB 2026. This is the author's version. The definitive Version of Record is available at [https://doi.org/10.1145/3807503.3819448](https://doi.org/10.1145/3807503.3819448)