发表机构
University of Piraeus(比雷埃夫斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明在LoRA微调中引入随机高斯条件噪声可提升组织病理学伪影检测的干净/伪影分离度,无需额外编码器,且效果经多轮消融和预注册实验验证。
AI 中文摘要
基于扩散的伪影检测器通过在干净组织上微调的模型,根据重建误差对全切片图像块进行评分。我们证明,将这种微调条件化为随机高斯嵌入——每一步从约200 KB的预计算嵌入统计中重新采样,无需编码器、无需缓存、无需改变推理——能持续扩大干净/伪影分离度。一个四步消融链表明,该益处既不需要内容(打乱的真实嵌入)、来源(合成高斯)、调谐强度(在8倍方差范围内平坦),也不需要逐块身份(每一步的新噪声);一个LoRA-dropout对照表明,是条件化通路本身,而非一般的权重扰动,承载了该效应。补丁级别的增益+0.25-0.48 Cohen's d在九次训练中重复出现;诚实的留一玻片评估在2/2个种子中通过了预注册标准;在281例数据集上的两个预注册外部端点确认了合并Delta F1 = +0.0073(95% CI)和+0.0129(97.5% CI,两次查看校正)。我们发布了完整的评估协议,包括测量的种子噪声和选择乐观定价。
英文摘要
Diffusion-based artifact detectors score whole-slide image patches by reconstruction error under a model fine-tuned on clean tissue. We show that conditioning this fine-tuning on random Gaussian embeddings -- resampled at every step from approx. 200 KB of precomputed embedding statistics, with no encoder, no cache, and no change to inference -- consistently widens the clean/artifact separation. A four-step ablation chain shows the benefit requires neither content (shuffled real embeddings), provenance (synthetic Gaussians), a tuned intensity (flat across an 8x variance range), nor per-patch identity (fresh per-step noise); a LoRA-dropout control shows the conditioning pathway specifically, not generic weight perturbation, carries the effect. Patch-level gains of +0.25-0.48 Cohen's d replicate across nine trainings; honest leave-one-slide-out evaluation clears a pre-registered bar in 2/2 seeds; and two pre-registered external endpoints on a 281-case set confirm pooled Delta F1 = +0.0073 (95% CI) and +0.0129 (97.5% CI, two-look corrected). We release the full evaluation protocol, including measured seed noise and selection-optimism pricing.
Comments23 pages, 3 figures. Code and laboratory record: https://github.com/kmouts/condnoise-histoqc (doi:10.5281/zenodo.22702198). Data: doi:10.5281/zenodo.22702800. Companion study: arXiv:2608.30835