arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04143cs.LGstat.ML

考虑电路:超越单一Softmax的深度分离与普适性

Consideration Circuits: Depth Separation and Universality Beyond a Single Softmax

Junjie Xiao, Huiwen Jia

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出考虑电路(CC),一种基于MNL单元有向无环图的多阶段选择模型,证明其深度3即可逼近任意选择表,且深度分离优于单一softmax,实验显示其参数少且性能更优。

中文摘要 AI 辅助

大多数基于特征的选择模型,无论是经典的还是深度的,都对项目进行评分并应用单一的softmax。我们引入了考虑电路(CC),这是一种基于特征的多阶段选择模型,由多项逻辑(MNL)单元的有向无环图定义。源单元为菜单项分配概率,内部单元使用从其概率加权特征摘要计算出的MNL权重来组合前驱分布。在具有固定非共线特征的三项折中任务中,菜单无关的随机效用模型(RUM),包括单个MNL单元,其误差被限制在远离零的范围内。相比之下,对于CC,我们建立了一个尖锐的深度-范数分离:将深度从2增加到3,对于误差ε的最优最大口味向量范数从Θ(log(1/ε)/ε)减少到Θ(log(1/ε))。深度2的下界对任意宽度和菜单无关的路由偏差都成立,而一个具有零路由偏差的五节点深度3电路达到了对数速率。更一般地,我们刻画了两个几何条件,这些条件是在有限菜单族上逼近任意确定性选择表的必要且充分条件。在这些条件下,深度3就足够了,而深度4在族包含非单例菜单时实现了最优的对数范数缩放。在实验中,具有少于600个参数的独立树电路在四个固定池基准和Expedia时间分割上,在评估模型中取得了最低的平均测试负对数似然(NLL)。作为输出头,CC推广了线性MNL读出,并在Expedia和Trivago上降低了每个测试编码器的平均测试NLL。

英文摘要

Most feature-based choice models, classical and deep, score items and apply a single softmax. We introduce consideration circuits (CC), feature-based models of multi-stage choice defined by directed acyclic graphs of multinomial logit (MNL) units. Source units assign probabilities to menu items, and internal units combine predecessor distributions using MNL weights computed from their probability-weighted feature summaries. On a three-item compromise task with fixed non-collinear features, menu-independent random-utility models (RUM), including a single MNL unit, suffer an error bounded away from zero. For CC, in contrast, we establish a sharp depth--norm separation: increasing depth from $2$ to $3$ reduces the optimal maximum taste-vector norm for error $ε$ from $Θ(\log(1/ε)/ε)$ to $Θ(\log(1/ε))$. The depth-$2$ lower bound holds for arbitrary width and menu-independent routing biases, while a five-node depth-$3$ circuit with zero routing biases attains the logarithmic rate. More generally, we characterize two geometric conditions that are necessary and sufficient for approximating arbitrary deterministic choice tables on finite menu families. Under these conditions, depth $3$ suffices, while depth $4$ achieves optimal logarithmic norm scaling whenever the family contains a non-singleton menu. In experiments, standalone tree circuits with fewer than $600$ parameters attain the lowest mean test negative log-likelihood (NLL) among the evaluated models on four fixed-pool benchmarks and the Expedia temporal split. As output heads, CC generalize the linear MNL readout and lower mean test NLL for every tested encoder on Expedia and Trivago.

发表机构

  • Peking University(北京大学)
  • University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑