arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于脑肿瘤视觉指令微调的多语言模型协作式MRI报告生成

Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology

Sinyoung Ra, Jonghun Kim, Hyunjin Park

arXiv 2607.14581首次发表:更新:

AI 中文总结

针对脑肿瘤缺乏配对3D成像 - 文本数据问题,提出用多语言模型协作创建3D图像 - 文本数据集,构建VLM将MRI扫描转换为令牌并与文本指令对齐,该方法在报告生成等任务中表现优于其他方法,有助于脑肿瘤诊断治疗。

AI 中文摘要

大语言模型(LLMs)及其向视觉语言模型(VLMs)的扩展,使文本与图像结合用于报告生成等任务变得更容易。医学领域现有的VLMs通常专注于二维图像(胸部X光),由于缺乏配对的三维成像 - 文本数据,将其扩展到三维成像一直很困难。因此,我们引入了一种新方法,利用胶质瘤和脑膜瘤病例的三维MRI扫描创建脑肿瘤的三维图像 - 文本数据集。我们使用一个协作系统,其中多个语言模型共同生成和检查报告,确保报告准确清晰。通过利用新的三维MRI - 文本数据集,我们进一步构建了一个VLM,将MRI扫描转换为令牌并使其与文本指令对齐。我们的VLM在报告生成和视觉问答任务中比其他二维和三维方法表现更好。我们的方法不仅提高了报告质量,还有助于脑肿瘤的更好诊断和治疗。

英文摘要

Recent advances in large language models (LLMs) and their extension to vision-language models (VLMs) have made it easier to combine text and images for tasks such as report generation. Existing VLMs in medicine typically focus on 2D images (chest X-rays), and their extension to 3D imaging has been difficult because of the lack of paired 3D imaging-text data. Thus, we introduce a new method for creating a 3D image-text dataset for brain oncology using 3D MRI scans of glioma and meningioma cases. We use a cooperative system in which several LLMs work together to generate and check reports, ensuring that they are accurate and clear. By leveraging the new 3D MRI-text dataset, we further build a VLM that converts MRI scans into tokens and aligns them with text instructions. Our VLM performed better in report generation and visual question answering tasks than other 2D and 3D methods. Our method not only improves the quality of reports but also helps with better diagnosis and treatment in brain oncology.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑