发表机构
ExtensityAI(ExtensityAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对从头编写ASP理论困难的问题,采用神经符号方法测试9种模型从LLM提炼ASP理论的能力,在VQA基准上验证了前沿模型的性能,并发布了相关资源。
AI 中文摘要
从头编写答案集编程(ASP)理论是一项困难且耗时的任务。我们采用神经符号方法,研究在固定包含求解器的智能体框架下,模型是否能提炼出完整且正确的理论。该协议与数据集无关:模型仅需一个提示和空文件作为起点,有1小时时间限制来推导完整理论。我们选择视觉问答(VQA)作为应用领域,三个基准测试(CLEVR、GQA、CLEVRER)因其公开可用且具有挑战性。为研究解决该任务所需的模型规模,我们测试了9种不同模型:4个前沿模型(Claude Sonnet 4.6、Claude Opus 4.7、GPT-5、DeepSeek V4 Pro)、2个中端模型(DeepSeek V4 Flash、gpt-oss-120b)和3个开源权重模型(qwen3.6-27b、gpt-oss-20b、qwen3.5-9b)。4个前沿模型中3个在CLEVR上达到100%,在GQA上达到92.8%-98.8%;在CLEVRER上,Sonnet、Opus、DeepSeek V4 Pro得分92.7%-95.3%。GPT-5在CLEVR上达到98.7%,但在GQA上降至41.8%,在CLEVRER上降至86.7%。添加其他数据集的手写参考理论,使其他三个前沿模型的准确率变化最多为±3.4个百分点,但使GPT-5的准确率下降3-19个百分点。我们发布了代码、提示和提炼出的理论。
英文摘要
Writing Answer Set Programming (ASP) theories from scratch is a difficult and time-consuming task. We take a neurosymbolic approach to study whether a model can distill complete and correct theories, given a fixed agent harness with the solver in the loop. The protocol is dataset-agnostic: with a single prompt and an empty file as the starting point the model is given a 1-hour time limit to derive a complete theory. We chose VQA as the application domain, three benchmarks (CLEVR, GQA, CLEVRER), as these are publicly available and non-trivial. In order to study the model scale required for solving this task we nine different models: four frontier (Claude Sonnet 4.6, Claude Opus 4.7, GPT-5, DeepSeek V4 Pro), two mid-tier (DeepSeek V4 Flash, gpt-oss-120b), and three open-weights (qwen3.6-27b, gpt-oss-20b, qwen3.5-9b). Three of four frontier models reach 100% on CLEVR and 92.8%-98.8% on GQA; on CLEVRER, Sonnet, Opus, DeepSeek V4 Pro score 92.7%-95.3%. GPT-5 reaches 98.7% on CLEVR but drops to 41.8% on GQA and to 86.7% on CLEVRER. Adding handwritten reference theories from other datasets moves the other three frontier models by at most +/-3.4 pp but reduces GPT-5's accuracy by 3-19 pp. We release the code, prompts, and theories distilled.
CommentsAccepted at NeSy 2026