发表机构
Johns Hopkins University; Columbia University; The Hong Kong Polytechnic University; The Pennsylvania State University(约翰霍普金斯大学; 哥伦比亚大学; 香港理工大学; 宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估Jev作为科学工作流中语义决策组件的性能,通过测试框架比较十二种模型配置在十个科学案例中的表现,发现Jev在完全语义正确性上与其他五种配置相当,且成功响应延迟最低,同时揭示了错误选择对下游计数的影响。
AI 中文摘要
科学工作流通常需要在确定性计算进行之前,从已知关系中进行选择。观测结果是否共享文化、治疗或参考标准,会改变由此产生的计数或比较的科学含义。我们使用一个遵循其文档化指导并将算术分配给代码的测试框架,评估 Jev 作为语义决策组件的性能。该研究在十个科学案例中,对二十个基于来源的选择,比较了十二种模型配置,每个案例重复五次。我们分别测量语义选择、下游输出和最终声明标签。Jev 在完全语义正确性方面与其他五种配置相匹配,并在成功响应中实现了最低的中位延迟。在三个比较模型中,一个文化历史问题上的七次错误选择改变了下游计数,同时保持了正确的最终标签。这些结果确定了 Jev 在准备好的科学决策任务中的有用角色,并说明了为什么评估该角色需要检查工作流将重复使用的关系和数量。
英文摘要
Scientific workflows often require choosing among known relations before a deterministic calculation can proceed. Whether observations share a culture, treatment or reference standard can change the scientific meaning of the resulting count or comparison. We evaluate Jev as a semantic decision component using a harness that follows its documented guidance and assigns arithmetic to code. The study compares twelve model configurations on twenty source-grounded Choices across ten scientific cases, each repeated five times. We measure semantic selections, downstream outputs and final claim labels separately. Jev matched five other configurations at complete semantic correctness and achieved the lowest observed median latency among successful responses. Across three comparison models, seven wrong selections on one culture-history question changed downstream counts while preserving the correct final label. These results identify a useful role for Jev in prepared scientific decision tasks and show why evaluating that role requires checking the relations and quantities that a workflow will reuse.
Comments10 pages, 1 figure, 5 tables. Includes references and appendices