发表机构
Obuda University; Biomatics and Applied Artificial Intelligence Institute, John von Neumann Faculty of Informatics, Obuda University; Physiological Controls Research Center, Obuda University; HUN-REN Centre for Agricultural Research; Czech Academy of Sciences(欧布达大学; 欧布达大学约翰·冯·诺依曼信息学院生物信息学与应用人工智能研究所; 欧布达大学生理控制研究中心; HUN-REN农业研究中心; 捷克科学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Distribird是一款自动化文献先验分布设计的智能体网络应用,通过多智能体流程生成贝叶斯模型校准所需先验,在24个跨领域参数上验证,其先验质量与单提示基线相当,且具备可追溯性、拒绝越权请求及本地运行的优势。
AI 中文摘要
基于过程的模型的贝叶斯校准需要为每个模型参数指定先验分布。尽管已有数十年的方法学研究,研究人员几乎总是退而使用均匀先验,主要原因是从科学文献构建信息先验的过程缓慢,且需要同时具备领域知识和统计专业技能。我们提出了**Distribird**,一款自动化该过程的智能体网络应用。给定参数名称、物理描述和领域上下文,Distribird会部署多智能体流程:搜索文献、提取并按领域相关性对报告值加权,再通过AIC模型选择拟合概率分布。若无可供文献,系统会采用合理的无信息先验作为替代,并明确报告所生成的每一个先验背后的依据及其置信水平。该工具适用于模型参数具有物理解释性、且已发表文献中存在领域知识的场景。我们在10个科学领域的24个参数上对该工具进行评估,将三个开源权重模型(Qwen3.6 27B、Gemma 4 31B、Mistral Small 4 119B)与单提示大语言模型基线进行对比。在先验质量上,完整流程与该基线表现相当;每一个先验都可追溯到其构建所依据的具体论文和数值;内置的有效性层会对超出范围的请求拒绝生成先验,而单提示基线在30个模型-参数案例中有11个会为超出范围的请求返回有信心但无依据的先验;此外,每一次大语言模型调用均在本地运行,因此不会将参数描述或未发表的建模细节传输给第三方大语言模型提供商(仅生成的搜索词会发送至公共文献数据库)。对于科学应用而言,我们认为这些特性比点估计精度的边际提升更为重要。
英文摘要
Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present Distribird, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, the system falls back to sensible uninformative alternatives, and clearly reports both the evidence behind and the confidence level of every prior it produces. It is designed for the problems where the models have physically interpretable parameters, where domain knowledge exists in the published literature. We evaluate the tool on 24~parameters across 10 scientific domains comparing three open-weight models (Qwen3.6 27B, Gemma 4 31B, Mistral Small 4 119B) with a single-prompt LLM baseline. On prior quality the full pipeline matches this baseline. Every prior is traced to the specific papers and values from which it was constructed; a built-in validity layer declines to produce priors for out-of-scope requests, whereas the single-prompt baseline returns confident but unfounded priors for them in 11 of 30~model-parameter cases; and every language-model call runs locally, so no parameter description or unpublished modelling detail is transmitted to a third-party LLM provider (only generated search terms reach the public literature databases). For scientific use, we argue these properties matter more than a marginal improvement in point-estimate accuracy.