通过自适应获取与顺序融合的高效多模态推理
Efficient Multimodal Inference through Adaptive Acquisition and Sequential Fusion
- Arm Inc.(Arm公司)
- Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SemARC通过自适应获取与顺序融合,在推理时动态选择模态并提前停止,在六个数据集上平均提升宏F1 3.2%,降低61.4%的GFLOPs和44%的延迟,实现高效多模态推理。
AI中文摘要:
多模态系统通常对所有可用输入进行编码,即使其中一部分输入已足以进行预测。自适应获取可以通过利用从增量融合证据中得到的预测来决定接下来应编码哪种模态以及何时停止,从而降低这一成本。然而,顺序融合使得这些预测依赖于顺序,因此基于这些预测的决策可能需要区分同一已获取集合的众多阶乘历史。我们提出了SemARC,它将顺序模态聚合器(SeMA)与自适应运行时控制器(ARC)相结合,并在每个模态的编码器运行之前利用已获取的证据来选择该模态。SeMA仅执行选中的编码器和融合分支,更新固定大小的状态,并在每次获取后无需重新计算先前分支即可进行预测。我们在随机化的模态子集和顺序下监督每个获取前缀,以鼓励跨获取顺序的一致预测。ARC结合了基于集合的边际效用先验与残差拟合Q学习,以选择下一个可用模态或停止,而无需检查未获取的输入或保留获取顺序。在六个多模态分类数据集和十一个基线上,相对于每个数据集上最准确的基线,SemARC平均实现了3.2%更高的宏F1分数和61.4%更低的总推理GFLOPs。端到端延迟在GPU和CPU上平均降低44.0%,在Android INT8上相对于最快的实测基线平均降低47.2%。在运行时模态缺失变化的情况下,SemARC仍会跳过可用模态,在24种条件中的21种条件下达到或超过最佳基线的宏F1分数,同时平均总GFLOPs降低14.8%。因此,SemARC为跨异构设备的高效多模态推理提供了一条实用路径。
英文摘要:
Multimodal systems often encode every available input, even when a subset suffices for prediction. Adaptive acquisition can reduce this cost by using predictions from incrementally fused evidence to decide which modality to encode next and when to stop. However, sequential fusion makes these predictions order-dependent, so decisions based on them may need to distinguish factorially many histories of the same acquired set. We introduce SemARC, which couples a Sequential Modality Aggregator (SeMA) with an Adaptive Runtime Controller (ARC) and uses acquired evidence to select each modality before its encoder runs. SeMA executes only selected encoder and fusion branches, updates a fixed-size state, and predicts after each acquisition without recomputing earlier branches. We supervise every acquisition prefix under randomized modality subsets and orders to encourage consistent predictions across acquisition orders. ARC combines a set-dependent marginal-utility prior with residual fitted-Q learning to select the next available modality or stop, without inspecting unacquired inputs or retaining acquisition order. Across six multimodal classification datasets and eleven baselines, SemARC achieves 3.2% higher macro-F1 and 61.4% lower total inference GFLOPs on average relative to each dataset's most accurate baseline. End-to-end latency falls by 44.0% across GPU and CPU and by 47.2% on Android INT8 relative to the fastest measured baseline, on average. Under varying runtime modality missingness, SemARC still skips available modalities, matching or exceeding the best baseline macro-F1 in 21 of 24 conditions with 14.8% lower total GFLOPs on average. SemARC thus offers a practical path toward efficient multimodal inference across heterogeneous devices.