来自40亿参数开放模型的随机试验校准答案:一项注册测试与许可证清洁发布
Calibrated Answers About Randomized Trials From a 4-Billion-Parameter Open Model: A Registered Test and a License-Clean Release
浏览论文内容
中文总结 AI 辅助
Fiorillo v0.5是一个40亿参数开放模型,通过注册标准验证其在随机试验证据推断上的校准性能,并以Apache许可证清洁发布。
中文摘要 AI 辅助
Fiorillo v0.5是一个开放模型,它对每个答案以概率形式回答类型化问题。其主要专家模块读取随机试验的文章(截断至6,144个标记),并回答干预措施与对照相比是否显著增加、显著减少或未显著改变结局(证据推断2.0,EI)。该模型基于Qwen3-4B-Base,配备低秩适配器和决策头,仅在2,657篇训练文章中许可证允许复用的1,431篇上进行EI微调。在本次版本测试预测之前,在开放科学框架上注册了四项标准以决定其发布,第二项标准在EI测试集上评估,该测试集的标签是公开的。在该测试集(333篇文章中的1,218个提示)上,预期校准误差为0.0168,而限值为0.05;对数损失比先验低0.8603(95%区间0.8104至0.9078),比读取相同输入的Gemma 4 31B-it低0.1829(0.1164至0.2598);宏F1为0.9248,而对比为0.8668,因此所有四项标准均通过。仅使用清洁文章训练相同配方在准确性上损失0.0123(0.0034至0.0207;描述性)。在没有文章的情况下,宏F1降至0.4384;仅标题将其提高了0.0939(0.0655至0.1234),这可能由陈述试验结果或回忆的标题解释;交换干预和对照使其方向答案反转了0.6652。按发布版本运行,文件在预先设定的限值内与评估预测匹配。该发布基于Apache License 2.0(数字对象标识符 https://doi.org/10.57967/hf/10722 )。
英文摘要
Fiorillo v0.5 is an open model that answers typed questions with a probability for each answer. Its main specialist reads a randomized trial's article, cut to 6,144 tokens, and answers whether an intervention significantly increased, significantly decreased or did not significantly change an outcome against a comparator (Evidence Inference 2.0, EI). It is Qwen3-4B-Base with low-rank adapters and a decision head, fine-tuned for EI only on the 1,431 of 2,657 training articles whose own license allows reuse. Four criteria registered on the Open Science Framework before this version's test predictions decided its release, the second bar judged on EI's test split, whose labels are public. On that split (1,218 prompts in 333 articles), the expected calibration error was 0.0168 against a limit of 0.05; log loss was below the prior's by 0.8603 (95 percent interval 0.8104 to 0.9078) and below that of Gemma 4 31B-it, reading the same input, by 0.1829 (0.1164 to 0.2598); and macro-F1 was 0.9248 against 0.8668, so all four criteria passed. Training the same recipe on clean articles alone cost 0.0123 in accuracy (0.0034 to 0.0207; descriptive). With no article, macro-F1 fell to 0.4384; the title alone raised it by 0.0939 (0.0655 to 0.1234), which a title stating the result or recall of the trial could explain; exchanging intervention and comparator reversed 0.6652 of its direction answers. Run as released, the files matched the evaluated predictions within limits set in advance. The release is under the Apache License 2.0 (digital object identifier 10.57967/hf/10722).