发表机构
Amazon.com, Inc.(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究首次实证评估Jev模型在网络流量分类中的表现,发现标记示例显著提升其准确率,但整体仍逊于传统树集成模型,且成本更低。
AI 中文摘要
我们仅使用前十个数据包的大小、方向和包间时间,在CESNET-QUICEXT-25数据集上对Jev进行了十个数据集定义的应用标签的评估。据我们所知,这是首次对通用决策模型(此处以Jev为代表)进行网络流应用分类的实证研究。在训练期后26个采集周的52,000条记录中,40个固定标记示例将Jev的准确率从9.80%提升至28.42%。在8,000条记录上训练的随机森林和极端随机树分别达到69.95%和66.80%的准确率,并且在每一周都优于Jev。将Jev的上下文增加到150个示例,在第一个测试周达到34.50%的准确率。在配对的100条记录子集上,使用40个示例的Jev在0.750秒的中位请求时间内达到29%的准确率,而通过Azure使用高推理努力的生成式语言模型OpenAI GPT-5.6 Sol则为37%和6.036秒;Jev还产生了较低的API费用。配对子集并未确立任一服务的准确性优势,且时间差异反映了不同的服务配置。因此,标记示例显著改善了Jev,但所测试的Jev配置仍不如训练的树集成准确;监督预算的不平等和固定配置使得无法将差距归因于单一原因。
英文摘要
We evaluate Jev on ten dataset-defined application labels in CESNET-QUICEXT-25 using only the first ten packets' sizes, directions, and inter-packet times. To the best of our knowledge, this is the first empirical study of general-purpose decision models, represented here by Jev, for application classification of network flows. Across 52,000 records from 26 collection weeks following the training period, 40 fixed labeled examples raise Jev's accuracy from 9.80% to 28.42%. Random Forest and Extra Trees trained on 8,000 records achieve 69.95% and 66.80% and outperform Jev in every week. Increasing Jev's context to 150 examples yields 34.50% on the first test week. On a paired 100-record subset, Jev with 40 examples achieves 29% accuracy at a median request time of 0.750 s, versus 37% and 6.036 s for the generative language model OpenAI GPT-5.6 Sol with high reasoning effort through Azure; Jev also incurs lower API charges. The paired subset does not establish an accuracy advantage for either service, and the timing reflects different service configurations. Thus, labeled examples substantially improve Jev, but the tested Jev configurations remain less accurate than trained tree ensembles; unequal supervision budgets and fixed configurations prevent attributing the gap to a single cause.