AI 中文总结
本文推导Dirichlet平滑转移模型下添加轨迹的边际对数似然增量,证明其为KL散度的加权缩减,并给出上界、协方差恒等式及交互符号判据;案例研究显示单个轨迹预测精度与批处理质量无必然联系。
AI 中文摘要
对于Dirichlet平滑的转移模型,向训练档案中添加一条工作流轨迹的效果是参考加权对数似然的一个精确变化。我们推导出这一变化,并证明它是参考条件分布与模型之间Kullback-Leibler散度的加权缩减。由此形式,我们获得了任何采集策略可获增益的上界,该上界将毫奈特差异表示为可实现增益的份额,一个关于参考加权效应的精确协方差恒等式,以及两个候选之间交互的符号判据,据此批处理目标既非子模也非超模。针对BPI Challenge 2012贷款申请日志的案例研究测量了这三项指标,并在参考加权与预算单位的四种组合之一中发现了正向选择结果。在该组合中,两个回归器拟合于相同的描述符和相同的标签,其中一个对单个增量的预测准确度远高于另一个,中位$R^2$分别为0.87和0.62,但实现的增益份额较小,分别为61%和69%,因此对单个轨迹的排序准确性对于批处理质量既非充分也非必要条件。
英文摘要
For a Dirichlet-smoothed transition model, the effect of adding one workflow trace to the training archive is an exact change in reference-weighted log likelihood. We derive that change and show that it is a weighted reduction of Kullback--Leibler divergence between the reference conditionals and the model. From this form we obtain an upper bound on the gain available to any acquisition, which expresses a millinat difference as a share of what is attainable, an exact covariance identity for the effect of the reference weighting, and a sign criterion for the interaction between two candidates, from which the batch objective is neither submodular nor supermodular. A case study on the BPI Challenge 2012 loan-application log measures all three and finds a positive selection result in one of the four combinations of reference weighting and budget unit. There, of two regressors fitted to identical descriptors and identical labels, the one that predicts individual increments far more accurately, median $R^2$ 0.87 against 0.62, realizes the smaller share of the attainable gain, 61 against 69 per cent, so ranking accuracy for individual traces is neither necessary nor sufficient for batch quality.
Comments14 pages, 1 figure, 3 tables