AI 中文总结
本文针对非纯净环境中AI系统的约束观测问题,提出结合平衡准确率与多惩罚项的自适应验证前沿表示选择器,在聚焦基准中可提升前沿分数并减少特征数量,但泛化性有限。
AI 中文摘要
部署在非纯净基准环境中的AI系统通常依赖不完整、不稳定、成本高或因监控故障而降级的观测数据。本文研究约束观测下的表示选择问题:当原始准确率并非唯一操作标准时,选择状态表示的方法。我们提出一种验证前沿选择器,将平衡准确率与特征成本、过拟合 gap、验证-测试不稳定性的惩罚项相结合。在使用三个 scikit-learn 数据集、五种观测 regime、45 个匹配任务单元、720 个候选动作和 405 个表示行的聚焦公共表格基准中,该自适应选择器相较于全轨迹特征,将前沿分数提高了 0.025801,同时将平均特征数量减少了 22.733。平衡准确率差异很小且无统计学意义。更广泛的离线压力测试结果不一。因此,支持的结论是有限的:自适应表示选择可在匹配的基准设置中改善约束观测下的鲁棒性-效率前沿,但并不能普遍优于轨迹基线。
英文摘要
AI systems deployed outside clean benchmark settings often rely on observations that are incomplete, unstable, costly, or degraded by monitoring failures. This paper studies representation selection under constrained observation: choosing a state representation when raw accuracy is not the only operational criterion. We propose a validation-frontier selector that combines balanced accuracy with penalties for feature cost, overfit gap, and validation-test instability. In a focused public-tabular benchmark using three scikit-learn datasets, five observation regimes, 45 matched task cells, 720 candidate actions, and 405 representation rows, the adaptive selector improves frontier score over full trace features by 0.025801 while reducing mean feature count by 22.733. Balanced-accuracy difference is small and not statistically significant. A broader offline stress test gives mixed results. The supported claim is therefore bounded: adaptive representation selection can improve a constrained-observation robustness-efficiency frontier in matched benchmark settings, but does not universally dominate trace baselines.