发表机构
Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对比5种推理时代理策略,发现多图像医学推理中,顺序投票决策规则比扩展进化搜索预算的影响更大,其最终测试准确率显著优于基线及复杂变体。
AI 中文摘要
多图像医学视觉问答(VQA)不仅是提示长度的问题,更是代理决策的根本挑战。医学视觉语言代理必须聚合有序图像间的证据,对答案顺序扰动保持鲁棒性,并避免过拟合至嘈杂的搜索时反馈。本研究通过在相同高预算ShinkaEvolve配置下优化的5种推理时代理策略的受控对比,对MedFrameQA展开研究,评估采用可复现的内部冻结划分(含1331个进化样本、665个保留样本、855个最终测试样本)。在5次独立重复运行中,最强方法为最简单的鲁棒聚合器:顺序投票策略实现57.89±0.65%的最终测试准确率,显著优于固定基线(52.73±0.42%)及更复杂但脆弱的顺序重排变体(55.79±0.43%),配对自助分析证实了这些显著增益。将进化搜索预算从50代扩展至100代未带来泛化收益:尽管保留样本性能略有提升,但最终测试准确率从57.89%降至56.02%。研究表明,对于多图像医学推理,定义正确的代理决策规则比扩大优化搜索预算的影响大得多。
英文摘要
Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents must aggregate evidence across ordered images, remain robust to answer-order perturbations, and avoid overfitting to noisy search-time feedback. We study MedFrameQA through a controlled comparison of five inference-time agentic strategies, optimized using the same high-budget ShinkaEvolve configuration and evaluated on a reproducible internal frozen split (1,331 evolution, 665 holdout, 855 final test). Across five independent repeated runs, the strongest method emerges as the simplest robust aggregator: the \textbf{order-vote} policy achieves $57.89 \pm 0.65\%$ final-test accuracy, significantly outperforming the fixed baseline ($52.73 \pm 0.42\%$) and the more complex, albeit brittle, order-rerank variant ($55.79 \pm 0.43\%$). Paired bootstrap analysis confirms these significant gains. Extending the evolutionary search budget from 50 to 100 generations yields no generalization benefit: while holdout performance marginally increases, final-test accuracy drops from $57.89\%$ to $56.02\%$. Our findings suggest that for multi-image medical reasoning, defining the correct agentic decision rule is substantially more impactful than expanding the optimization search budget.
CommentsPresented at the CVPR 2026 Workshop on Multi-Modal Reasoning for Agentic Intelligence