arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RamseyGadgets:面向大语言模型(LLMs)的图构造数据集

RamseyGadgets: A Graph Construction Dataset for LLMs

Zohair Raza Hassan, Deepak Pandita

arXiv 2608.14999首次发表:更新:

AI 中文总结

本研究提出新型图构造数据集RamseyGadgets,评估5个开源LLMs在其70个拉姆齐良图问题上的表现,发现LLMs在困难问题上准确率仅37.70%,Gemma-4-31B性能最优,还可用于确定提升LLMs表现的提示类型。

AI 中文摘要

构造特殊图是图论与计算机科学中的重要任务,许多流行的图构造成果源于对相关图的全面探索与人类创造力。鉴于生成式AI在数学领域的应用日益广泛,自然会想到测试LLMs是否能运用推理能力构造出具有指定属性的图。遗憾的是,许多自然图构造问题(如寻找极值拉姆齐良图,即避免特定单色子图)已在文献中被广泛研究,因此难以确定某一构造是LLM推理能力的产物还是其从训练数据中回忆得到的。在本研究中,我们提出了RamseyGadgets——一个包含70个未被充分探索的图构造问题的新型数据集,这些问题要求寻找具有特殊属性的拉姆齐良图(例如包含一条固定颜色的边)。这些问题的解规模适中(最多10个顶点),可通过SAT求解器验证,适合自动评估。我们的数据集易于扩展,只需更改需避免的单色子图即可得到一组新问题。我们在该数据集上评估了5个开源LLMs的性能并报告结果。研究发现,LLMs在数据集的困难层级问题上仅达到37.70%的准确率,其中Gemma-4-31B在5个模型中性能最高。我们还展示了该数据集如何帮助确定何种提示能帮助LLMs在该任务上表现更好。

英文摘要

Constructing special graphs is an important task within graph theory and computer science. Many popular graph constructions are the result of a comprehensive exploration of relevant graphs and human ingenuity. Given the rise of generative AI usage in mathematics, it is natural to test whether LLMs are able to construct graphs with specified properties using their reasoning capabilities. Unfortunately, many natural graph construction problems, such as finding extremal Ramsey-good graphs (i.e., avoiding specific monochromatic subgraphs), have been explored extensively in the literature, making it difficult to ascertain whether a construction is the product of an LLM's reasoning capabilities or its recollection from training data. In this work, we introduce \textbf{RamseyGadgets}, a novel dataset of 70 underexplored graph construction problems that require finding Ramsey-good graphs with special properties (e.g., containing an edge with a fixed color). These problems have reasonably sized solutions (at most 10 vertices) that can be verified by SAT solvers, making them suitable for automatic evaluation. Our dataset is easily expandable, as one can simply change the monochromatic subgraphs being avoided to obtain a new set of problems. We evaluate the performance of five open-source LLMs on our dataset and report the results. Our findings show that LLMs achieve only 37.70% accuracy on the hard-tier problems in our dataset, with Gemma-4-31B achieving the highest performance out of the five. We also showcase how our dataset allows us to ascertain what kind of hints help LLMs perform better at this task.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑