发表机构
Tokyo Institute of Science High School; National Institute of Advanced Industrial Science and Technology (AIST); Visual Geometry Group, University of Oxford(东京理科大学附属高中; 国立先进工业科学技术研究所; 牛津大学视觉几何组)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态语言模型能否评估书法局部笔触质量并生成反馈,构建评估框架让三个模型评估图像对并与专家打分比较,还研究了RAG变体,结果显示模型有一定绝对分数准确性,但与专家排名相关性不强,RAG有正负两方面表现。
AI 中文摘要
本文研究多模态语言模型能否评估书法中的局部笔触质量并生成具有教育意义的自然语言反馈。构建了一个评估框架,其中三个多模态语言模型(GPT-4o、Claude Sonnet 4和Gemini 2.5 Flash)使用五点序数尺度评估书法作品的前后图像对,并将它们的输出与三位专家书法家给出的分数进行比较。还研究了Claude的检索增强生成(RAG)变体作为初步条件。结果表明所有模型都达到了有用水平的绝对分数准确性(MAE),GPT-4o表现最佳(MAE = 0.885)。然而,没有一个模型与人类专家产生统计学上显著的总体排名相关性(肯德尔tau)。对生成理由的词汇分析揭示了每个模型的特征性评估偏差,RAG显示出提高排名相关性但降低绝对准确性,这对基于文本的规则注入来说是一个重要的负面结果。
英文摘要
This paper investigates whether multimodal LLMs can evaluate local brushstroke quality in calligraphy and generate educationally useful natural language feedback. We construct an evaluation framework in which three multimodal LLMs (GPT-4o, Claude Sonnet 4, and Gemini 2.5 Flash) assess before-after image pairs of calligraphic works using a five-point ordinal scale, and compare their outputs against scores assigned by three expert calligraphers. We additionally examine a Retrieval-Augmented Generation (RAG) variant of Claude as a preliminary condition. Results show that all models achieve useful levels of absolute score accuracy (MAE), with GPT-4o performing best (MAE = 0.885). However, none of the models produce statistically significant overall rank correlations with human experts (Kendall's tau). Vocabulary analysis of generated rationales reveals characteristic evaluative biases in each model, and RAG is shown to improve rank correlation while worsening absolute accuracy, constituting an important negative result for text-based rule injection.
Comments6 pages, 8 figures. Accepted to the SAUAFG Workshop at CVPR 2026