发表机构
Universitat Politècnica de Catalunya (UPC); Universitat Pompeu Fabra (UPF); International Center for Numerical Methods in Engineering (CIMNE)(加泰罗尼亚理工大学; 庞培法布拉大学; 国际工程数值方法中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对科学文献数学表达式自动翻译受高质量数据集稀缺阻碍的问题,提出开源可定制框架MioFFAn,基于MioGatto架构扩展功能,通过大语言模型实现部分自动化,用标准NLP指标评估,初步证明人机协作方法有效。
AI 中文摘要
科学文献中数学表达式到可执行符号代码的自动翻译(公式形式化)受高质量技术科学领域真实数据集稀缺的阻碍。本文提出MioFFAn,一个开源、以文档为中心且可定制的框架,用于促进此任务的快速注释。基于MioGatto架构扩展功能,克服结构限制并引入公式形式化特定功能。允许用户配置自定义分类法等,确保框架适用于不同科学领域。还通过大语言模型纳入部分自动化,定义模块化子任务,用标准NLP指标评估策略,进行了初步评估证明其人机协作方法的有效性。
英文摘要
The automatic translation of mathematical expressions in scientific literature into executable symbolic code (a process we refer to as Formula Formalization) is hindered by a severe scarcity of high-quality, ground-truth datasets specialized for technical scientific domains. In this paper, we present MioFFAn, an open-source, document-centric, and customizable framework designed to facilitate rapid annotation for this task. Building upon the MioGatto architecture, we extend existing features to overcome structural limitations and pivot its scope by introducing specific functionalities for Formula Formalization, such as selection of equations of interest and aided symbolic code specification. By allowing users to configure custom taxonomies and properties for identified symbols, and compatible symbolic operators, we ensure the framework is adaptable to diverse specialized scientific fields. Furthermore, MioFFAn is designed to incorporate partial automation via Large Language Models. By defining a modular set of automated sub-tasks with strict output formats, we enable researchers to iteratively refine automation capabilities and evaluate competing strategies using standard NLP metrics. We specify the current automation methodology and perform a preliminary evaluation that demonstrates to efficacy of this human-in-the-loop approach.
CommentsPresented in the 3rd International Workshop on Natural Scientific Language Processing (NSLP 2026), co-located at LREC2026
Journal refProceedings of the 3rd Int. Workshop on Natural Scientific Language Processing (NSLP 2026) at LREC 2026, pages 206-217