AI 中文总结
本文提出Tactus模型,仅用低成本压力阵列数据实现开放词汇对象识别,在STAG基准上表现优于或匹配闭集CNN,小数据方案及传感器校准是其关键,相关资源已公开。
AI 中文摘要
电阻式压力阵列是成本最低、应用最广的触觉传感器,但触觉表征学习却集中在对变形凝胶成像的光学传感器上。本文提出Tactus,一种仅从压力数据回答文本查询的开放模型:在STAG基准(27个物体、预留的录音数据)上,它在四次运行中达到0.771±0.062的top-1准确率(top-3为0.935),与该数据集的有监督闭集CNN(0.76)相当,且在未训练分类器头的情况下达到最优。该方法适用于小数据:187条训练录音,在14.4万个未标记的同传感器帧上进行掩码自编码器预训练,以及传感器自身的仿射校准,其恢复的准确率超过所有架构调整的总和。发布的模型错误集中在少数接触模糊类,与文本目标几何无相关性(702个类对的斯皮尔曼rho≤0.05),在改述甚至仅用名称的查询中,准确率仍保持在1个点内;两个不同帧可恢复8帧准确率的89%。失败案例以同等精度报告:跨传感器预训练池化无增益,视觉协同训练会降低触觉性能,输入管道的错误归一化在生成合理中间结果的同时,悄悄丢弃了传感器97%的动态范围。模型权重、代码及模型接入的内存层均已公开发布。
英文摘要
Resistive pressure arrays are the cheapest and most widely shipped tactile sensors, yet tactile representation learning has concentrated on optical sensors that image a deforming gel. We present Tactus, an open model that answers text queries from pressure data alone: on the STAG benchmark (27 objects, held-out recordings), it reaches 0.771 +/- 0.062 top-1 over four runs (top-3 0.935), matching, and at best exceeding, the dataset's supervised closed-set CNN at 0.76, with no trained classifier head. The recipe is small-data: 187 training recordings, masked-autoencoder pretraining on 144k unlabeled same-sensor frames, and the sensor's own calibration affine, which recovered more accuracy than every architecture change combined. The released model's errors concentrate in a few contact-ambiguous classes, are uncorrelated with text-target geometry (Spearman rho <= 0.05 over 702 class pairs), and survive paraphrased and even bare-name queries within one point; two diverse frames recover 89% of eight-frame accuracy. Failures are reported with equal precision: cross-sensor pretraining pooling gave no gain, vision co-training degraded touch, and a mis-normalized input pipeline silently discarded 97% of the sensor's dynamic range while producing plausible intermediate results. Weights, code, and the memory layer the model plugs into are released openly.
Comments6 pages, 4 figures. Weights: https://huggingface.co/EximiusLabs/fusion-embedding-2-tactus