AI 中文总结
该研究针对超高维特征选择问题,提出随机序列搜索(SSS)框架,将序列方法的全量遍历替换为带预算的采样步算子,其衍生的sSFFS在多个高维数据集上优于现有排序方法,计算成本与维度无关。
AI 中文摘要
序列子集搜索——带浮动回溯的前向选择及其衍生方法——仍是特征选择的质量基准,但该家族的每个成员在每一步都会遍历全部剩余候选特征池,这使其无法应用于超高维问题;在超高维问题中,仅保留单个特征排序方法实用,而这类方法对特征交互的建模很弱或完全不建模。我们引入了一种带预算的采样步算子对,用每一步固定数量的候选评估替代全量遍历。候选通过温度控制的softmax采样从每一步搜索执行的所有准则评估中在线学习的、依赖特征的单特征统计量中抽取,并有均匀探索下限保障;每一步的计算成本与维度无关。替换这些算子可将任何序列方法转化为其随机对应版本,定义了随机序列搜索(SSS)家族;我们研究了浮动搜索的随机对应版本sSFFS。在500维的madelon数据集上,sSFFS在每个子集大小下都保留了至少97%的全量SFFS准则值,而评估量仅为其四分之一左右,同时相同预算下的均匀采样在madelon的协同特征上表现崩溃。在5000维的gisette数据集上,远超全量SFFS处理范围的情况下,sSFFS在匹配预算下超过了DAF和BIF排序的饱和准则水平;留一验证显示,在500个训练样本时,超出序列前沿的约束是准则而非搜索。在10105维的reuters数据集上,在可信的多项式过滤准则下,sSFFS在搜索目标和留一准确率上,在每个子集大小下都优于BIF和DAF,单核评估仅需约两分钟。本文附带了一个经过验证的独立实现。
英文摘要
Sequential subset search -- forward selection with floating backtracking and its descendants -- remains the quality reference in feature selection, but every member of the family sweeps the full pool of remaining candidate features at each step, which excludes it from very-high-dimensional problems; there, only individual-feature ranking remains practical, and it models feature interplay weakly or not at all. We introduce a budgeted sampled step operator pair that replaces the full sweeps by a fixed number of candidate evaluations per step. Candidates are drawn by temperature-controlled softmax sampling from dependency-aware per-feature statistics learned online from every criterion evaluation the search performs, guarded by a uniform exploration floor; per-step cost becomes independent of dimensionality. Substituting the operators turns any sequential method into its stochastic counterpart, defining the Stochastic Sequential Search (SSS) family; we study the stochastic counterpart of floating search, sSFFS. On 500-dimensional madelon, sSFFS retains at least 97% of the full-SFFS criterion value at every subset size at about a quarter of its evaluations, while uniform sampling at the same budget collapses on madelon's synergistic features. On 5,000-dimensional gisette, far beyond full-SFFS reach, sSFFS exceeds the saturated criterion level of DAF and BIF ranking at matched budgets; holdout validation shows that at 500 training samples the binding constraint beyond the sequential frontier becomes the criterion, not the search. On 10,105-dimensional reuters, under a trustworthy multinomial filter criterion, sSFFS dominates BIF and DAF on the search objective and on holdout accuracy at every subset size, in about two minutes of single-core evaluation work. A verified standalone implementation accompanies the paper.
Comments23 pages, 5 figures, 2 tables. Ancillary files: a standalone implementation of sSFFS (MIT license, verified bit-identical to the reference implementation), a checksum-verified converter building the exact experiment datasets from the UCI originals, the derived Reuters-21578 term matrix used in the experiments (research use only), and the configuration files of all experiments