潜在有序证据与未对齐输出:多模态大语言模型的推理时有序视角对齐
Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs
- Monash University(莫纳什大学)
- Faculty of Information Technology, Monash University(莫纳什大学信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态大语言模型有序输出未对齐问题,本文提出推理时的Ordinal Lens Alignment方法,提升有序回归任务性能且优于现有基线。
AI中文摘要:
多模态大语言模型(Multimodal LLMs)将语言模型接口应用于视觉输入,其中年龄估计、图像质量评估、疾病分级等有序回归任务需要基于有序类别标签的自回归决策。本文探究多模态大语言模型是否能可靠地将内部有序证据转换为有序数字标记输出。在四个有序基准和四个多模态大语言模型主干上,隐藏状态中的有序标签可线性恢复,斯皮尔曼相关系数最高达0.938,且任务设计的提示词进一步强化了该结构。然而原生数字标记输出仅微弱暴露该结构:未嵌入矩阵过滤了有序方向,在全部16种模型-数据集组合中,数字标记行空间保留率低于1.15%,线性探针与原生输出的准确率存在16至77个绝对点的差距。本文提出Ordinal Lens Alignment(OLA,有序视角对齐),这是一种冻结主干的推理时方法,在中深层解码器层训练轻量的W_S锚定视角,将其融合为有序分布,仅在生成时校正数字标记的对数几率。该方法在多数设置下优于SOTA LoRA微调的OrderChain基线,同时保持多模态大语言模型冻结状态,在多数单元中优于判别式有序基线,且在所有设置下均优于离线视角。
英文摘要:
Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, and disease grading require autoregressive decisions over ordered class labels. We ask whether MLLMs reliably convert internal ordinal evidence into ordered digit-token outputs. Across four ordinal benchmarks and four MLLM backbones, ordinal labels are linearly recoverable from hidden states with Spearman correlation up to 0.938, and a task-designed prompt further sharpens this structure. Yet native digit-token outputs weakly expose it: the unembedding matrix filters the ordinal direction, and the digit-token row space retains below 1.15% across all 16 model-dataset combinations, with a 16 to 77 absolute-point accuracy gap between linear-probe and native outputs. We introduce Ordinal Lens Alignment (OLA), a frozen-backbone inference-time method that trains lightweight W_S-anchored lenses on mid-to-deep decoder layers, fuses them into an ordinal distribution, and corrects only digit-token logits at generation. OLA outperforms the SOTA LoRA-tuned OrderChain baseline in most settings while keeping the MLLM frozen, surpasses discriminative ordinal baselines in most cells, and improves over an offline lens in every setting.