LLM2Jev:LLM 已是 Jev 风格决策模型——何时以及如何微调它们
LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
查看机构详情
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出 LLM2Jev 框架,利用 LLM 的下一词元概率直接输出 Jev 风格决策,免训练即可匹配社区模型,微调仅对弱模型和特定任务有针对性提升,并保持对话生成不退化。
中文摘要 AI 辅助
Jev 风格决策模型返回预定义选项上的分类概率分布,而不生成自由形式的文本,从而使软件系统能够直接基于其输出采取行动。在这项工作中,我们研究了通用 LLM 在多大程度上已经天然具备这种能力,以及何时真正需要微调。我们提出了 LLM2Jev,一个保持架构的框架,直接从括号化数字标识符的下一个词元概率中提取校准的决策。LLM2Jev 既提供免训练推理方案,也提供微调目标,该目标通过树因子化列表损失优化候选选择,同时使用 KL 散度惩罚将辅助预测锚定到基础模型。在 Qwen3.5-4B 和 Qwen3-0.6B 上的评估中,我们发现现代 LLM 本质上是有效的决策模型:无需训练,4B 模型即可匹配基于相同主干构建的社区 Jev 风格模型,优于字母逻辑读取,支持任意数量的选项,并原生处理图像上的多模态决策。微调提供的是针对性而非普遍性的收益——显著改进较弱的模型和特定任务(如多选项意图路由),但对强主干模型的收益递减。关键的是,我们的 KL 锚点防止了对话文本生成中的行为退化,LoRA 在能力强的模型上提供了最强的性能。
英文摘要
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM-as-Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.