arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PetQA:兽医知识与临床推理基准测试

PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning

Taegyun Kim, Youngwook Ham, Jungwook Rhim, Ju-Hyun An, Sungkyu Park, Kunwoo Park

arXiv 2609.04598首次发表:更新:

发表机构

Soongsil University; Kangwon National University; KDI School of Public Policy and Management(崇实大学; 江原大学; KDI公共政策与管理学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出PetQA这一韩语兽医知识与临床推理基准测试,含万余对问答及多模态数据,评估18种模型在三种设置下的表现,提供多语言版本,助力开发可靠兽医AI系统。

AI 中文摘要

我们推出PetQA,这是一个用于评估大型语言模型(LLMs)和大型视觉语言模型(LVLMs)兽医知识与临床推理能力的韩语长文本问答(QA)基准测试。PetQA包含10076个纯文本问答对和8751个多模态问答对,这些问答对来源于关于犬和猫的真实问题,并配有兽医专家提供的答案。其测试集PetQA-Bench还包含问题类型和临床状况的标注。我们在三种设置(零样本推理、检索增强生成(RAG)、监督微调(SFT))下,使用ROUGE、BERTScore和LLM作为评判者的指标,对18个模型的事实性和有用性进行评估。基准测试结果概述了当前模型在处理兽医临床查询方面的优势与局限性,并强调需要更有效的适配方法来开发用于兽医护理的临床可靠AI系统。为便于更广泛使用,我们还提供了PetQA-Bench的五种语言翻译版本。

英文摘要

We introduce PetQA, a Korean long-form question-answering (QA) benchmark for evaluating veterinary knowledge and clinical reasoning in large language models (LLMs) and large vision-language models (LVLMs). PetQA contains 10,076 text-only and 8,751 multimodal QA pairs derived from real-world questions about dogs and cats, paired with answers from expert veterinarians. Its test split, PetQA-Bench, further includes annotations for question types and clinical conditions. We evaluate eighteen models using ROUGE, BERTScore, and LLM-as-a-judge metrics for factuality and helpfulness under three settings: zero-shot inference, retrieval-augmented generation (RAG), and supervised fine-tuning (SFT). The benchmarking results provide an overview of the strengths and limitations of current models in addressing veterinary clinical queries and highlight the need for more effective adaptation methods to develop clinically reliable AI systems for veterinary care. To facilitate broader use, we additionally provide translated versions of PetQA-Bench in five languages.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑