arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向边缘智能数据模型分类的小型语言模型:一种成本感知的混合方法

Small Language Models for Smart Data Model Classification at the Edge: A Cost-Aware Hybrid Approach

Cristian Martella, Angelo Martella, Antonella Longo, Motaz Saad

arXiv 2610.07093首次发表:更新:

发表机构

University of Salento(萨伦托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估轻量级开源语言模型在资源受限边缘环境下对智能数据模型分类的性能,系统对比多种架构与近零成本基线,为模型选择与部署提供优化准确性和效率的实用见解。

AI 中文摘要

物联网(IoT)中异构数据源的快速激增,涉及智慧城市、能源管理和环境监测等领域,亟需高效且可扩展的数据标准化方法。智能数据模型(SDM)的有效分类对于促进互操作性至关重要。然而,现有方法往往受限于高资源消耗,且在计算能力受限的边缘环境中缺乏适用性。为弥合这一差距,本研究评估了轻量级开源语言模型(LM)在资源受限条件下,将输入数据实体解析为对应最佳拟合SDM表示的性能。研究系统性地对多种模型进行了基准测试,包括通用型(GP)、推理专用型(RS)和代码专用型(CS)架构,并跨多个领域特定数据集进行验证。针对文献中缺乏轻量级、资源高效解决方案的现状,本研究为模型选择、任务制定和部署策略提供了重要且有价值的见解,以优化准确性和效率。一项补充实验还将所调查的大型语言模型(LLM)与两个近零成本相似性基线(词频-逆文档频率(TF-IDF)和轻量级句子编码器)在同一任务上进行比较,为解释基于LLM的分类在边缘平台上的实际价值提供了强有力的参考点。

英文摘要

The rapid proliferation of heterogeneous data sources within the Internet of Things (IoT) across domains such as smart cities, energy management, and environmental monitoring necessitates efficient and scalable data standardization methods. Effective classification of smart data models (SDMs) is essential for facilitating interoperability. However, existing approaches are often limited by high resource consumption and lack applicability in edge environments with constrained computational capabilities. Aiming to bridge this gap, the proposed study evaluates the performance of lightweight open-source language models (LMs) to resolve an input data entity against its corresponding best fitting SDM representation under resource-constrained conditions. It systematically benchmarks a diverse array of models, including general purpose (GP), reasoning-specialized (RS), and code-specialized (CS) architectures, across multiple domain-specific datasets. Addressing the current omission of lightweight, resource-efficient solutions in the literature, the investigation provides significant and valuable insights into model selection, task formulation, and deployment strategies that optimize accuracy and efficiency. A complementary experiment also compares the surveyed large language models (LLMs) against two near-zero-cost similarity baselines (Term Frequency-Inverse Document Frequency (TF-IDF) and a lightweight sentence encoder) on the same task, providing a strong reference point for interpreting the practical value of LLM-based classification on edge platforms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑