arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33561cs.LG

GraphSelect:多模态图推断中的预算表示选择

GraphSelect for Budgeted Representation Selection in Multimodal Graph Inference

  • Shandong University(山东大学)
  • Beijing Institute of Technology(北京理工大学)
  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li

中文总结 AI 辅助

针对多模态图推断中的预算表示选择问题,提出GraphSelect方法,通过联合评估交换优化子集,在保留20%表示时仅损失0.10%准确率。

中文摘要 AI 辅助

多模态图预测器结合文本、图像和关系来对连接的实体进行分类。为了保持其预测,需要多少这样的输入?我们研究了预算表示选择,该任务在每种模态的独立容量限制下,从候选文本和图像向量中选择一个子集。完整候选输入的预测定义了需要保留的类别。挑战在于,一个表示的贡献取决于其他选中的输入,而图传播将其影响扩展到多个节点。我们的实证研究表明,候选排名会随选中的输入而变化,而在类别停止变化后,预测概率仍然具有信息量。更新分数可以改善选择,而交换输入可以改进一个容量已满的子集。这些发现催生了GraphSelect,它从单个候选增益开始,通过联合评估的交换来优化子集。它筛选有前景的移除和添加,在接受能减少预测损失的交换时更新分数。在六个图上的实验显示,其平均目标恢复率高于六种为选择任务调整的归因和解释方法。在两个图的九个训练架构上,每种模态保留20%的候选表示,相对于完整候选输入,平均准确率下降0.10个百分点,从而以显著更少的文本和图像表示保持了分类性能。

英文摘要

Multimodal graph predictors combine text, images, and relations to classify connected entities. How much of this input is needed to preserve their predictions? We study budgeted representation selection, which chooses a subset of candidate text and image vectors under a separate capacity for each modality. Predictions from the complete candidate input define the classes to preserve. The challenge is that a representation's contribution depends on the other selected inputs, while graph propagation extends its effects across nodes. Our empirical study shows that candidate rankings change with the selected input, while predicted probabilities remain informative after the class stops changing. Updating scores improves selection, and exchanging inputs can improve a subset whose capacity is already filled. These findings lead to GraphSelect, which starts from individual candidate gains and refines the subset through jointly evaluated exchanges. It screens promising removals and additions, accepts an exchange when it reduces the prediction loss, and updates the scores. Experiments on six graphs show higher mean objective recovery than six attribution and explanation methods adapted to the selection task. Across nine trained architectures on two graphs, retaining 20% of the candidate representations per modality gives a mean accuracy drop of 0.10 percentage points relative to full candidate input, preserving classification performance with substantially fewer text and image representations.

↑