发表机构
The Hong Kong Polytechnic University; The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文建立三阶段研究框架,在不更新多模态蛋白质语言模型参数的前提下,通过优化推理阶段采样策略,揭示默认推理协议的次优性并提升其任务性能,得出与现有共识不同的基础模型相关结论。
AI 中文摘要
多模态蛋白质语言模型(pLMs)学习蛋白质序列-结构的联合分布,其生成性能也应关键依赖于推理阶段的采样策略。然而,现有研究更多关注模型训练,而非推理阶段策略的表现。本文中,我们建立了一个三阶段研究框架,针对三个代表性多模态pLMs和四项基础任务,对其推理设计空间开展实证研究。我们在多模态pLMs上评估了普通采样、任务特定的无分类器引导采样,以及奖励引导的束搜索,分别对应对采样分布、每一步对数几率、并行轨迹的控制。围绕探索-利用权衡的互补改进,我们(1)揭示了默认推理协议的次优性,确定了面向任务的采样偏好;(2)观察到各任务上的显著定量增益,在不更新模型参数的情况下持续提升多模态pLMs的上限性能;(3)得出了与现有共识不同的关于基础模型的结论。
英文摘要
Multimodal protein language models (pLMs) learn joint protein sequence-structure distributions, and their generation performance should also depend critically on inference-time sampling strategies. Yet prior work has focused more on model training than on how inference-time strategies behave. In this paper, we establish a three-stage investigation framework to empirically study the inference design space of multimodal pLMs across three representative pLMs and four fundamental tasks. We evaluate vanilla sampling, task-specific classifier-free guidance, and reward-guided beam search on multimodal pLMs, corresponding to controls over sampling distributions, per-step logits, and parallel trajectories. Throughout the complementary advancements centered on exploration-exploitation trade-off, we (1) reveal the suboptimality of default inference protocols and identify task-oriented sampling preferences; (2) observe substantial quantitative gains across tasks, consistently boosting the upper bound performance of multimodal pLMs without updating model parameters; (3) derive conclusions about base models that differ from prior consensus.
CommentsAccepted to EMNLP 2026 Main Conference