arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

轻量级多模态模型能否评估LLM的推理性能?一项面向计算最优文档推理的研究

Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference

Zishan Ahmad, Vishal Vaddina

arXiv 2608.18591首次发表:更新:

发表机构

Phi Labs; Quantiphi(菲实验室; 宽梯斐科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出轻量级多模态模型DRB,基于新基准BudgetDoc实现文档任务中LLM推理预算的动态分配,在降低成本的同时多数情况下达到或优于固定最大预算的性能,具备跨模型泛化潜力。

AI 中文摘要

为LLM统一分配推理预算成本高昂且易产生过度推理惩罚,尤其在视觉布局驱动复杂性的文档任务中更是如此。为解决该问题,我们推出BudgetDoc——首个多模态基准,可为三类文档任务的模型-预算-性能权衡提供明确监督。利用BudgetDoc,我们训练了DRB(Document-Reasoning Balancer,文档推理平衡器),这是一个约10亿参数的预评估器(SigLIP-2 + Qwen3-0.6B),可预测不同预算水平下的序数模型性能,达到0.753的加权F1值。在为五个前沿模型和三个数据集动态分配推理预算时,DRB在15种配置中有9种达到或优于始终采用最大预算的基线模型的F1值,同时大幅降低了成本。最后,初步评估表明DRB具备向跨模型选择泛化的潜力。

英文摘要

Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking penalties; especially in document tasks where visual layouts drive complexity. To address this, we introduce BudgetDoc, the first multimodal benchmark providing explicit supervision for model-budget-performance trade-offs across three document tasks. Using BudgetDoc, we train DRB (Document-Reasoning Balancer), an approx. 1B-parameter pre-flight estimator (SigLIP-2 + Qwen3-0.6B) that predicts ordinal model performance across budget levels, achieving a 0.753 weighted F1. When dynamically allocating reasoning budgets across five frontier models and three datasets, DRB matches or improves F1 scores compared to always-maximum-budget baselines in 9 of 15 configurations while drastically reducing cost. Finally, preliminary evaluations demonstrate DRB's potential to generalize to cross-model selection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑