arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-12-12 至 2025-12-12 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 2 篇

2506.11375 2025-12-12 cs.AI cs.CL 57%

Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables

在化学表格上评估多模态大语言模型的基准测试

Yitong Zhou, Mingyue Cheng, Qingyang Mao, Yucong Luo, Qi Liu, Yupeng Li, Xiaohan Zhang, Deguang Liu, Xin Li, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) University of Science and Technology of China(中国科学技术大学) Artificial Intelligence Research Institute(人工智能研究院) iFLYTEK Co., Ltd(iFLYTEK公司)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.AI

AI总结 本文提出ChemTable基准,用于评估多模态模型在理解化学表格中的能力,揭示了现有模型在跨模态对齐和领域推理方面的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10440 2025-12-12 cs.CL 50%

Enhancing Next-Generation Language Models with Knowledge Graphs: Extending Claude, Mistral IA, and GPT-4 via KG-BERT

通过知识图谱增强下一代语言模型:通过KG-BERT扩展Claude、Mistral IA和GPT-4

Nour El Houda Ben Chaabene, Hamza Hammami

机构 * STIH Laboratory, Sorbonne University(索邦大学STIH实验室) National Engineering School of Tunis(突尼斯国家工程学院) Faculty of Sciences of Tunis(突尼斯科学学院)

专题命中 视觉问答 :grounding(abstract)

AI总结 通过KG-BERT将知识图谱与Claude、Mistral IA和GPT-4结合,提升其事实可靠性与上下文感知能力。

Comments This paper was accepted and scheduled for inclusion in the ICALT 2025 proceedings but was ultimately not published due to absence from the conference presentation. It appears in the official program booklet. Conference: 2025 IEEE International Conference on Advanced Learning Technologies (ICALT)

详情

展开后加载摘要…

URL PDF HTML 收藏