发表机构
RIKEN Center for Biosystems Dynamics Research; Institute of Medicine, University of Tsukuba; Laboratory Automation Suppliers' Association(理化学研究所生物动态研究中心; 筑波大学医学研究院; 实验室自动化供应商协会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建了八项期刊编辑任务的基准,在紧凑型工作站上评估20个本地开放权重模型,发现性能与权重规模无单调关联,结合确定性检查器可覆盖全部违规,表明此类工作站可胜任编辑工作。
AI 中文摘要
期刊开始考虑使用语言模型处理稿件,但提交的稿件尚未发表,且若政策禁止将其发送至外部服务,模型必须在期刊控制的硬件上运行。本地托管模型在编辑工作中的能力尚未被测量。在此,我们根据期刊的《作者须知》构建了一个包含八项编辑任务的基准测试,任务来源于我们植入并独立验证缺陷的稿件,以及一篇预印本的已发表评审意见,并在实验室或小型编辑办公室可采用的紧凑型桌面工作站上,评估了二十个开放权重模型,其权重规模跨度达二十五倍。最强的模型检测出40个植入的指南违规中的36个,占用81 GB内存;一个17 GB的模型检测出33个。在我们测试的最佳配置中,我们未观察到权重规模与得分之间一致的单向关联:每项任务的秩相关性可忽略不计(Spearman |rho| <= 0.19),且在同一模型家族内,较大的成员得分低于其较小的兄弟模型。一个不使用模型的正则表达式和算术确定性检查器在不到一秒内检测出相同的违规中的31个,其检测结果与最强模型的检测结果合并后覆盖了全部40个。在单一的同行评审案例中,最佳模型从三篇已发表评审中恢复了12分中的6分。提示结构在单个模型内大幅改变了得分。因此,一旦编写了确定性检查器,此类工作站的性能水平即可达到。
英文摘要
Journals are beginning to consider language models for manuscript handling, but submitted manuscripts are unpublished, and where policy forbids sending them to an external service the model must run on hardware the journal controls. The capability of locally hosted models on editorial work has not been measured. Here we constructed a benchmark of eight editorial tasks from a journal's Instructions for Authors, from manuscripts carrying defects we seeded and verified independently, and from published reviews of a preprint, and evaluated twenty open-weight models spanning a twenty-five-fold range of weight size on a compact desktop workstation of the kind a laboratory or small editorial office can adopt. The strongest model detected 36 of 40 seeded guideline violations and occupied 81 GB; a 17 GB model detected 33. Across the best configurations tested, we observed no consistent monotonic association between weight size and score: rank correlations were negligible on every task (Spearman |rho| <= 0.19), and within one model family the larger member scored below its smaller sibling. A deterministic checker of regular expressions and arithmetic, using no model, detected 31 of the same violations in a fraction of a second, and the union of its detections with those of the strongest model covered all 40. On the single peer-review case, the best model recovered 6 of 12 points from three published reviews. Prompt structure substantially altered scores within individual models. This level of performance is therefore within reach of a workstation of this class, once the deterministic checks are written.