arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在大小、延迟和隐私约束下的设备端商业意图检索:一个具有类型化出口边界的 3 MiB 检索系统

On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A 3 MiB Retrieval System with Typed Egress Boundaries

Hyojung Han

arXiv 2610.00170首次发表:更新:

AI 中文总结

本研究在设备端大小、延迟和隐私约束下构建 3 MiB 商业意图检索系统,基于 6,020 叶子分类法,发现锚点缺失导致精度下降,教师模型可弥补容量差距,并验证锚点有效性。

AI 中文摘要

我们研究完全在用户设备上运行的商业意图推断,该推断受三个在工作开始前就已确定的约束条件限制:下载的负载小于 3 MiB,Tier-0 推断在 p95 下小于 20 ms,且不允许原始文本、内容嵌入或稳定标识符离开设备。在这些约束下,我们构建了一个基于 6,020 个叶子节点的商业分类法的检索路径:一个从韩语句子转换器蒸馏得到的静态嵌入表,量化为 4 位,无推断运行时。我们的主要结果是该约束在何处造成精度损失。在由他人标注的真实韩语商业文本(22,900 条 AI-Hub 购物评论)上,真实产品名称的中类别 top-5 准确率为 75.0%,而排列基线为 18.4%,但根据一个可观察特征进行划分:包含某个叶子名称作为子串的查询得分为 83.5%,不包含任何叶子名称的查询得分为 45.2%。一个规模为 196.6 倍的通用教师模型似乎定位了该差距(有锚点时 +20.1 个百分点,无锚点时 +0.1 个百分点)。该零效应是两个效应相互抵消的结果:同一个教师模型在学生自己的对比对上进行微调后达到 0.8586,在有锚点的情况下比纯编码器学生高 +10.6 个百分点,在无锚点的情况下高 +20.9 个百分点。成本并非均匀分布,但并非在任何地方都是免费的;在无锚点的情况下,任务适应对教师模型毫无帮助,因此受约束编码器在该处缺乏的是容量。高成本区域可在设备端通过排序器自身的得分边际检测到:放弃最不自信的五分之一查询可将其余查询提升至 0.8296。我们首次报告的第二个轴,即制造商型号代码,在源类别固定效应下不显著(-4.0 个百分点,p=0.51);锚点则显著(+13.0 个百分点)。负载为 2,942,652 字节,所有三个库链接均已测量。Tier-0 p95 在两部 iPhone(A14、A16)上分别为 4.431 和 3.670 ms,在预算型 Android 平板(Snapdragon 695)上为 5.080 ms,均慢于运行相同代码的三台服务器 CPU。分类法监督主要来自合成韩语话语。

英文摘要

We study commercial intent inference that runs entirely on the user's device, under three constraints frozen before the work began: the downloaded payload under 3 MiB, Tier-0 inference under 20 ms at p95, and no raw text, content embedding, or stable identifier leaving the device. Under them we build a retrieval path over a 6,020-leaf commercial taxonomy: a static embedding table distilled from a Korean sentence transformer, quantized to 4 bits, no inference runtime. Our main result is where that constraint costs accuracy. On real Korean commerce text labelled by others (22,900 AI-Hub shopping reviews), mid-category top-5 on real product names is 75.0% against an 18.4% permutation baseline, but splits on one observable: a query containing some leaf name as a substring scores 83.5%, one containing none 45.2%. A generic 196.6x larger teacher seemed to localize the gap (+20.1 pp without an anchor, +0.1 with). That null was two effects cancelling: the same teacher fine-tuned on the student's own contrastive pairs reaches 0.8586 and beats the pure-encoder student by +10.6 pp with an anchor and +20.9 pp without. The cost is not uniform, but it is not free anywhere; where the anchor is absent, task adaptation buys the teacher nothing, so what the constrained encoder lacks there is capacity. The expensive regime is detectable on-device from the ranker's own score margin: declining the least confident fifth lifts the rest to 0.8296. A second axis we first reported, a manufacturer model code, does not survive source-category fixed effects (-4.0 pp, p=0.51); the anchor does (+13.0 pp). Payload is 2,942,652 bytes, all three library links measured. Tier-0 p95 is 4.431 and 3.670 ms on two iPhones (A14, A16) and 5.080 ms on a budget Android tablet (Snapdragon 695), all slower than three server CPUs on the same code. Taxonomy supervision is mostly synthetic Korean utterances.

Comments36 pages, 40 tables, 1 figure. Evidence bundle (measurement ledgers and verification gates, no datasets): https://github.com/hyojunguy/ondevice-intent-evidence

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑