AI 中文总结
LAVOIR通过将缺失信息槽位置于输入中,利用摊销信息价值指导单遍决策编码器何时及询问什么,无需人工标注,在受控和真实对话中显著提升准确率。
AI 中文摘要
“系统一”决策模型,如TypeSafe的Jev及其开源对应模型Laya,能够在单次前向传播中以校准概率回答关于文本的带类型问题,但它们无法请求缺失信息:当第一条消息未说明两个部门之间的区别时,它们会进行猜测。我们提出LAVOIR(Laya与信息价值路由),它将候选的缺失信息片段(槽位)放置在输入中紧邻答案选项的位置,从而使得一次前向传播既能返回决策分布,又能针对每个槽位返回如果询问用户该信息时正确决策概率的预期增益。信息价值目标无需人工标注:黄金决策来自模式规则,大型语言模型仅负责将消息和答案进行语言化表达,另一家族的模型检查每个文本,并且将每条消息与多个配置文件配对,使得对已实现增益的回归能够估计预期增益。基尼不纯度上限将预测值限制在校准模型仍能获得的增益范围内。在一项受控研究中,在已见模式上的决策与贝叶斯上限在统计上无显著差异。最终模型的提问策略在已见模式上与贪婪预言机信息价值策略相匹配(AUC 0.799对比0.797),并且每次对话最多提问0.5个问题,比从不提问准确率高14.1个百分点。在真实的ABCD对话中,一次真实交流在LAVOIR提问的情况下将准确率提高8.3个百分点,而在不提问的情况下保持不变;在SGD上,上限将提问率从93%降至8.6%。在Laya的十二个基准测试中,LAVOIR在七个基准上的得分高于Laya报告的分数,并且回答一个问题仅需31毫秒(中位数,GH200)。
英文摘要
"System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what separates two departments, they guess. We present LAVOIR (Laya with Value-Of-Information Routing), which places the candidate pieces of missing information (slots) in the input next to the answer options, so that one forward pass returns both the decision distribution and, for every slot, the expected gain in the probability of the correct decision if the user were asked about it. VOI targets need no human labels: gold decisions come from schema rules, an LLM only verbalizes messages and answers, a model from another family checks every text, and pairing each message with several profiles makes regression on realized gains estimate the expected gain. A Gini-impurity cap bounds the predicted value by what a calibrated model can still gain. In a controlled study, decisions on seen schemas are statistically indistinguishable from the Bayes ceiling. The final model's question policy matches a greedy oracle VOI policy on seen schemas (AUC 0.799 vs. 0.797), and with at most 0.5 questions per conversation it is 14.1 points more accurate than never asking. On real ABCD conversations, one real exchange raises accuracy by 8.3 points where LAVOIR asks and leaves it unchanged where it does not; on SGD the cap lowers the asking rate from 93% to 8.6%. On Laya's twelve benchmarks LAVOIR is above Laya's reported scores on seven, and it answers a question in 31 ms (median, GH200).
Comments11 pages, 3 figures, 7 tables. Code: https://github.com/moganai/lavoir ; model: https://huggingface.co/moganai/lavoir