AI 中文总结
本研究提出SBERT2S1将生物医学检索编码器转换为类型化决策模型,并构建BIODECIDE套件和MEDLINE-S1数据集,发现检索预训练对PFR头有益,但RLCD配方因奖励归一化导致性能落后,交叉熵校准更优。
AI 中文摘要
类型化决策模型在一次前向传播中对文本回答受模式约束的问题,并返回用于阈值化的概率。我们探究了为检索而训练的生物医学句子编码器是否适合作为此类模型的起点。我们提出了SBERT2S1,它将Sentence-Transformers编码器转换为双编码器、交叉头(C)和先验融合残差(PFR)决策模型,同时构建了BIODECIDE,一个生物医学类型化决策套件,以及MEDLINE-S1,包含从NLM索引中导出的243k个训练决策。在六对父检索器组合中,检索训练提高了零样本匹配内容承载选项的能力。微调后,其效果取决于头部:在五对组合和三种训练集规模下,检索训练显著帮助了保留检索先验的PFR,在15次比较中10次有益,但对C仅1次有益、5次有害。对两个头部和五种训练目标的匹配网格显示,C在每个目标下均优于PFR,且已发布的开放系统一模型的RLCD配方在交叉熵上落后2.5-3.0个百分点。该差距主要源于其奖励归一化,该归一化将噪声得分函数项放大了3.6-15倍;无偏留一估计器弥补了大部分差距。经过温度缩放后,没有哪个目标在校准上明显优于交叉熵。我们发布了代码、MEDLINE-S1标签和一个模型。
英文摘要
Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head (C) and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing. Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-set sizes, retrieval training significantly helps PFR, which keeps the retrieval prior, in 10 of 15 comparisons, but helps C in one and hurts it in five. A matched grid of two heads and five training objectives shows that C outperforms PFR under every objective, and that the released RLCD recipe of open System One models trails cross-entropy by 2.5-3.0 points. The deficit stems mainly from its reward normalisation, which inflates the noisy score-function term 3.6-15-fold; an unbiased leave-one-out estimator recovers most of the gap. After temperature scaling, no objective is clearly better calibrated than cross-entropy. We release the code, the MEDLINE-S1 labels and a model.