CoSTAR:面向遗留系统现代化的数据合成驱动的约束感知COBOL节级摘要
CoSTAR: Data Synthesis-Driven Constraint-Aware COBOL Section Summarization for Legacy System Modernization
浏览论文内容
中文总结 AI 辅助
CoSTAR通过执行验证的数据合成和约束感知训练,解决COBOL节级摘要的数据稀缺与迁移约束问题,使7B/8B小模型在ROUGE-L等指标上平均提升25.38%以上,并超越Qwen3-235B。
中文摘要 AI 辅助
COBOL对于政府、金融机构和大型企业仍然至关重要;然而,老化的技术、日益减少的专业知识和缺失的文档使得基于COBOL的遗留系统的现代化变得日益紧迫。在迁移之前,代码摘要是一种支持遗留系统理解的常见做法。然而,COBOL代码摘要,尤其是在节级层面,面临两个关键挑战:数据稀缺和迁移约束的保留。为了解决这些挑战,我们提出了CoSTAR,一个将执行验证的数据合成与约束感知的模型训练相结合的集成框架。CoSTAR通过基于LLM的生成,重新利用通用编程任务来合成执行验证的COBOL代码-摘要数据,以克服数据稀缺问题。基于合成数据,CoSTAR用相关的数据声明和自然语言解释来增强目标节,并使用约束引导的结构化理由来训练较小的基础LLM。训练后的LLM在COBOL节级摘要中保留了迁移约束。我们在公开的和机密的企事业COBOL系统上评估了CoSTAR。CoSTAR有效地合成了3,764个执行验证的训练实例。基于这些实例,建立在7B/8B基础LLM上的CoSTAR可以平均相对提升这些LLM在ROUGE-L上25.38%、在METEOR上53.84%、在chrF上37.22%的性能。在真实世界的企事业评估中,仅基于Qwen3-8B构建的CoSTAR在准确性、完整性和简洁性方面优于企事业部署的Qwen3-235B。这些结果表明,CoSTAR使小型、可本地部署的LLM在隐私敏感的COBOL遗留系统中能够达到与显著更大的LLM相当的性能。
英文摘要
COBOL remains critical to governments, financial institutions, and large enterprises; yet, aging technologies, shrinking expertise, and missing documentation make modernization of COBOL-based legacy systems increasingly urgent. Before migration, code summarization is a common practice to support legacy system understanding. However, COBOL code summarization, especially on section-level, faces two key challenges: data scarcity and migration constraint preservation. To address these challenges, we propose CoSTAR, an integrated framework that combines execution-validated data synthesis with constraint-aware model training. CoSTAR repurposes general-purpose programming tasks to synthesize execution-validated COBOL code-summary data through LLM-based generation to overcome data scarcity. Based on the synthesized data, CoSTAR augments target sections with relevant data declarations and natural-language explanations, and uses constraint-guided structured rationales to train smaller base LLMs. The trained LLMs preserve the migration constraints for COBOL section summarization. We evaluate CoSTAR on both public and confidential enterprise COBOL systems. CoSTAR effectively synthesizes 3,764 execution-validated training instances. Based on these instances, CoSTAR built on 7B/8B base LLMs can improve these LLMs with average relative gains of 25.38% on ROUGE-L, 53.84% on METEOR, and 37.22% on chrF. In real-world enterprise evaluation, CoSTAR built on only Qwen3-8B, outperforms the enterprise-deployed Qwen3-235B in accuracy, completeness, and conciseness. These results show that CoSTAR enables small, locally deployable LLMs to achieve performance competitive with substantially larger LLMs for privacy-sensitive COBOL legacy systems.
发表机构
- School of Software, Dalian University of Technology(大连理工大学软件学院)
- DUT Artificial Intelligence Institute(大连理工大学人工智能研究院)
- Hi-Think Technology, Corp.(海思科技)
机构由 AI 辅助整理,请以论文原文为准。