PainterBench:用于工具使用语言模型的图形发散思维基准
PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models
- University of Oulu(奥卢大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
PainterBench通过未完成图形任务评估工具使用语言模型的图形发散思维,测试了14个多模态模型,发现GPT-6 Astra最具创造力,并提供了自动化评分器ViDrA-adapted。
AI中文摘要:
图形发散思维是指将给定的形状片段发展成原创图画的能力。在人类中,这种能力通过未完成图形任务来评估。我们引入了PainterBench,一个将未完成图形任务移植到智能体环境的基准。智能体通过工具调用在画布上作画,并在每轮之后观察结果。画布包含一个无法擦除的起始形状,智能体的目标是将该形状融入其能产生的最具原创性的图画中。该任务是开放式的,智能体自行决定何时完成绘画。该基准测试了短视界内的增量视觉规划以及从预训练到多轮工具使用的创造力迁移能力。我们评估了从小型到前沿规模的14个多模态语言模型。在主要研究和六项敏感性分析中,我们收集了2,700幅图画,并为每幅图画以及300幅人类参考图画众包了创造力和可识别性评分。我们还提出了ViDrA-adapted,一个自动化评分器,可预测智能体图画的人类创造力评分(在随机保留测试集上r = 0.85)。图形发散思维在14个模型中差异很大,GPT-6 Astra产生了最具创造力的图画。相对于人类图画,智能体图画的创造力得分更高,但可识别性得分较低。我们发布了最终图画、每轮画布快照、工具调用轨迹、刺激库、基准测试工具、众包评分(N = 72,000)以及ViDrA检查点。
英文摘要:
Figural divergent thinking is the ability to develop a given shape fragment into an original drawing. In humans, this ability is assessed with incomplete-drawing tasks. We introduce PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting. The agent draws on a canvas through tool calls and observes the result after every turn. The canvas includes a starting shape which cannot be erased, and the agent's goal is to incorporate this shape into the most original drawing it can produce. The task is open-ended, and the agent itself decides when the drawing is finished. The benchmark tests incremental visual planning over a short horizon and the transfer of creative ability from pretraining to multi-turn tool use. We evaluate 14 multimodal language models from small to frontier scale. Across the primary study and six sensitivity analyses, we collect 2,700 drawings and crowdsource creativity and recognizability ratings for every drawing and for 300 human reference drawings. We also present ViDrA-adapted, an automated scorer that predicts human creativity ratings of agent drawings (r = 0.85 on random held-out test split). Figural divergent thinking varies widely across the 14 models, and GPT-6 Astra produces the most creative drawings. Relative to the human drawings, the agent drawings score higher in creativity but lower in recognizability. We release the final drawings, per-round canvas snapshots, tool call traces, stimulus bank, benchmark harness, crowdsourced ratings (N = 72,000), and ViDrA checkpoint.