arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00015cs.AI

结合检索增强生成流程的大语言模型优化与约束建模

Optimization and Constraint Modeling using LLMs with a Retrieval Augmented Generation Process

Prateek Roy, Akash Singirikonda

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出结合检索增强生成的大语言模型优化建模方法,通过合成数据集与检索指导提升Qwen 3 30B Instruct的优化建模准确率,为低成本部署LLM优化工具提供可行方案。

中文摘要 AI 辅助

优化建模与约束建模均为非平凡问题,需深厚的领域专业知识及建模形式语言熟练度。尽管它们在物流、医疗、供应链管理等领域至关重要,但当前大语言模型(LLM)常生成结构不一致或不完整的优化公式,尤其在组合场景中。本文评估基于精心构建的合成数据集搭建的检索增强生成(RAG)流程能否显著提升LLM的优化建模性能。研究人员使用Text2Zinc数据集的种子描述及由LLM生成的专业角色(以JSON格式指定并关联已验证的Python求解器脚本)合成了共500个优化问题,将这些问题编码至Chroma向量数据库中。对于每个推理问题,检索语义相似的问题并将其作为上下文指导LangChain LLM智能体。研究采用三个基准测试平台,在Qwen 3 30B Instruct模型下评估所提流程:NL4OPT上的准确率从40%升至72%,MAMO Easy上从40%升至56%,MAMO Complex上从32%升至56%。使用经语义验证的合成示例可大幅提升解决方案的准确率与结构合理性。合成数据集生成与检索增强的结合为微调提供了有效替代方案,表明领域特定合成语料库搭配检索增强可成为在实际决策支持场景中部署基于LLM的优化工具的实用途径,无需昂贵的模型再训练。

英文摘要

Both optimization modeling and constraint modeling are non-trivial problems requiring deep domain expertise and proficiency in modeling formalism languages. Despite their importance across logistics, healthcare, and supply chain management, current large language models regularly produce structurally inconsistent or incomplete optimization formulations, particularly in combinatorial settings. This paper evaluates whether a Retrieval-Augmented Generation pipeline built on a curated synthetic dataset can meaningfully improve LLM optimization modeling performance. A total of 500 optimization problems were synthesized using seed descriptions from the Text2Zinc dataset and professional personas created using an LLM, specified in JSON and associated with validated Python solver scripts. These problems were encoded in a Chroma vector database. For each inference problem, semantically similar problems were retrieved and used as contextual guidance for a LangChain LLM agent. Three benchmark testbeds were used to evaluate the proposed pipeline under the Qwen 3 30B Instruct model. Accuracy rose from 40% to 72% on NL4OPT, 40% to 56% on MAMO Easy, and 32% to 56% on MAMO Complex. The use of semantically validated synthetic examples greatly improves both solution accuracy and structure. The combination of synthetic dataset generation with retrieval augmentation provides an effective alternative to fine-tuning, suggesting that domain-specific synthetic corpora paired with retrieval augmentation can serve as a practical pathway for deploying LLM-based optimization tools in real-world decision-support contexts without costly model retraining.

补充信息

↑