arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用多模态大语言模型探索程序流程图在代码生成中的潜力

Exploring the Potential of Program Flowcharts on Code Generation Using Multimodal LLMs

Yuki Toi, Tao Xiao, Kazushi Tomoto, Masanari Kondo, Yasutaka Kamei

arXiv 2607.09146首次发表:更新:

AI 中文总结

研究利用多模态大语言模型探索程序流程图对代码生成的潜力,通过从示例代码生成流程图并与问题陈述一起提供给模型进行代码生成实验,发现结合流程图可提升性能,不同细节程度的流程图效果有差异,还比较了与少样本学习的效果。

AI 中文摘要

近年来,大语言模型取得显著进展,出现了能处理图像和音频等多种输入的多模态大语言模型。先前研究表明,提供文本和视觉信息可提升自动代码生成能力。在软件开发中,流程图等图表被广泛用于辅助代码理解。虽有研究探讨视觉输入对大语言模型及软件图表使用的影响,但流程图对多模态大语言模型性能的潜在影响仍未充分研究。本研究从AtCoder问题的示例解决方案代码生成流程图,并将其与问题陈述一起提供给GPT-4o用于代码生成。结果表明,将流程图与问题陈述相结合可使性能提高多达10%。此外,使用抽象流程图时,流程图细节程度增加与性能提升相关。还比较了提供流程图与少样本学习方法的有效性。结果表明,单样本学习能持续改进,而双样本学习仅有微小改进。我们的工作凸显了软件图表在支持多模态大语言模型驱动的代码生成中的重要性。

英文摘要

In recent years, Large Language Models (LLMs) have made significant strides, leading to the emergence of multimodal LLMs capable of processing diverse inputs such as images and audio. Previous research indicates that the supply of multimodal LLMs with combined textual and visual information improves the automatic code generation capabilities. In software development, diagrams such as flowcharts are widely employed to facilitate tasks like code comprehension. While existing studies investigated the impact of visual inputs on LLMs and the usage of software diagrams, the potential influence of providing flowcharts on multimodal LLM performance remains underexplored. In this study, we generated flowcharts from example solution code for AtCoder problems and provided these visual aids alongside problem statements to GPT-4o for code generation. Our findings demonstrate that integrating flowcharts with problem statements yields performance improvements of up to 10%. Furthermore, when employing abstracted flowcharts, we observed a trend indicating that increasing levels of flowchart detail correlate with enhanced performance. Additionally, we compared the effectiveness of flowchart provision to Few-Shot Learning approaches. The findings suggest that one-shot learning provides sustainable improvements, whereas two-shot learning results in only minor improvements. Our work highlights the importance of software diagrams in supporting multimodal LLM-driven code generation.

Comments21 pages, Accepted at the 26th IEEE International Conference on Software Quality, Reliability, and Security (QRS 2026), Regular Papers Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑