ChitraMiti:孟加拉语几何推理中的视觉定位与模态依赖基准
ChitraMiti: Benchmarking Visual Grounding and Modality Reliance in Bengali Geometric Reasoning
浏览论文内容
中文总结 AI 辅助
本文提出ChitraMiti-12.8k和NCTB-500基准,通过三阶段协议评估VLM在孟加拉语几何推理中的表现,发现结构化描述可替代图表,但模型跨模态验证能力不足,微调可提升性能。
中文摘要 AI 辅助
针对低资源语言以及需要同时阅读图表和问题的几何问题,视觉语言模型(VLM)的多模态数学推理评估仍然有限。我们引入了ChitraMiti-12.8k,一个包含12,874个孟加拉语平面几何问题及其结构化15属性描述的合成基准,以及NCTB-500,一个从孟加拉语教科书中手动提取的500个图表的补充集。通过一个三阶段协议,分别使用仅图表、图表加描述和仅描述输入,我们在五个开放权重和闭源VLM上表明,仅描述的性能在统计上与图表加描述的性能无显著差异,从而将结构化描述确立为受控评估的充分文本代理。尽管如此,模型在跨模态验证方面仍然表现不佳,即使它们正确回答了未修改的项目,也经常被交换的空间关系误导。我们进一步评估了在ChitraMiti-12.8k上的监督适应,发现微调提高了在ChitraMiti-1k和NCTB-500上的性能,尽管与最强的零样本模型相比仍有很大差距。总之,ChitraMiti-12.8k、NCTB-500和我们的评估协议为研究孟加拉语多模态几何推理提供了一种标准化方式,更广泛地说,为研究VLM是否真正根据所见检查其文本提供了一种标准化方式。我们的数据集和代码可在Hugging Face上公开获取,网址为https://this https URL。
英文摘要
Evaluation of vision-language models (VLMs) for multimodal mathematical reasoning remains limited for low-resource languages and for geometry problems that require reading a diagram and a question together. We introduce ChitraMiti-12.8k, a synthetic benchmark of 12,874 Bengali planar geometry problems paired with structured 15-attribute descriptions, and NCTB-500, a complementary set of 500 diagrams manually extracted from Bengali school textbooks. Using a three-phase protocol that separates diagram-only, diagram-plus-description, and description-only inputs, we show across five open-weight and closed-source VLMs that description-only performance is statistically indistinguishable from diagram-plus-description performance, establishing structured descriptions as a sufficient textual proxy for controlled evaluation. Despite this, models remain poor at cross-modal verification, frequently misled by a swapped spatial relation even when they answer the unmodified item correctly. We further evaluate supervised adaptation on ChitraMiti-12.8k, finding that fine-tuning improves performance on both ChitraMiti-1k and NCTB-500, although a substantial gap to the strongest zero-shot model remains. Together, ChitraMiti-12.8k, NCTB-500, and our evaluation protocol offer a standardized way to study Bengali multimodal geometry reasoning and, more broadly, whether VLMs actually check their text against what they see. Our dataset and code are publicly available on Hugging Face at https://huggingface.co/datasets/RaiyanKhaan/ChitraMiti.
发表机构
- North South University(北南大学)
机构由 AI 辅助整理,请以论文原文为准。