发表机构
Eindhoven University of Technology(埃因霍温理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SimAuthor提出一个持久创作框架,通过分布比较和结构化反馈迭代修订科学模拟器,在六个生物医学任务上优于基线,并提升泛化能力。
AI 中文摘要
基础模型能够生成科学代码,但创作一个科学模拟器(一个可执行程序,编码关于机制如何产生可观测信号的假设)需要迭代细化。科学充分性很少允许唯一的实现或精确的测试,因此模拟器必须根据有限的真实观测来判断。我们将这一场景研究为在弱经验反馈下的科学模拟器创作,其中模拟信号与真实信号之间的分布比较指导修订,目标本身是模拟器而不仅仅是其生成的样本。我们引入了SimAuthor,一个持久的创作框架,它保留并修订可执行的模拟器,将标量搜索分数与结构化差异反馈分开,并积累可复用的实现机制。我们在六个生物医学任务上评估了SimAuthor,涵盖心脏和呼吸音频、光电容积脉搏波(PPG)和心电图(ECG)。在固定的100次尝试预算下,SimAuthor在所有六个任务上优于PUCT分数搜索,通常优于文本策略优化,并在六个任务中的五个上取得了最高的端点分数。所创作的模拟器在未见过的录音上也有所改进,可迁移到独立的预训练表示,并在下游ECG分类中产生了显著的分布外增益。最后,138次审计修订中的111次改变了程序结构,占签署分数改进的86.1%。这些结果表明,持久的修订可以逐步将基础模型的知识转化为更好的可执行科学模拟器,基于有限的实证证据。
英文摘要
Foundation models can generate scientific code, but authoring a scientific simulator (an executable program encoding hypotheses about how mechanisms generate observable signals) requires iterative refinement. Scientific adequacy rarely admits a unique implementation or exact test, so simulators must instead be judged against limited real observations. We study this setting as scientific simulator authoring under weak empirical feedback, where distributional comparisons between simulated and real signals guide revision, and the target is the simulator itself rather than only its generated samples. We introduce SimAuthor, a persistent authoring harness that retains and revises executable simulators, separates scalar search scores from structured discrepancy feedback, and accumulates reusable implementation mechanisms. We evaluate SimAuthor on six biomedical tasks spanning cardiac and respiratory audio, photoplethysmography (PPG), and electrocardiography (ECG). Under a fixed 100-attempt budget, SimAuthor outperforms PUCT score search on all six tasks, generally outperforms textual-strategy optimization, and achieves the highest endpoint score on five of six. The authored simulators also improve on unseen recordings, transfer to independent pretrained representations, and yield substantial out-of-distribution gains in downstream ECG classification. Finally, 111 of 138 audited revisions alter program structure and account for 86.1% of the signed score improvement. These results suggest that persistent revision can progressively convert foundation-model knowledge into better executable scientific simulators from limited empirical evidence.