一图胜千词:视觉语言模型如何在提升准确率的同时降低AI能源成本
A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy
浏览论文内容
中文总结 AI 辅助
该研究提出将时间序列编码为二维图的视觉语言模型,可大幅减少输入词元、降低推理能耗,同时在电信异常检测等任务中提升准确率,解决了LLM处理多维度KPI数据的低效问题。
中文摘要 AI 辅助
大语言模型(LLM)推理占AI运营能源的90%以上,且与输入词元数直接成正比,这对电信网络分析和数值时间序列分析(NTSDA)而言是关键的低效问题——4G/5G基站的原始多维度关键绩效指标(KPI)窗口会扩展为数千个浮点词元。视觉语言模型(VLM)通过将时间序列编码为二维图消除了这种不匹配,在Llama-3.2-90B、Qwen2.5-VL-72B和Pixtral-12B架构上实现了3.6至10.4倍的输入词元减少。这转化为1.8至2.5倍的实测推理能源降低,在每15分钟间隔监控200个基站的电信边缘部署和云无线接入网(CloudRAN)中,每天可节省约7.2兆焦耳(MJ)。关键的是,效率提升并未牺牲准确率:微调后的Llama-3.2-90B-Vision VLM的精度比仅文本的对应模型高220.7%,在电信异常检测上比LSTM和ARIMA基线模型高出144%以上。在公共基准测试中,Pixtral-12B的J/F1分数提升了20.6倍,平均F1值为0.82。在24个KPI下,文本表示超出了大多数生产级LLM的128K上下文窗口,导致仅文本处理若不截断则不可行,而视觉表示仍在标准范围内。这些结果确立了VLM作为数值时间序列工作负载的节能且准确率更优的模态,为将能源消耗作为首要工程约束的AI推理系统提供了实证依据。
英文摘要
LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network analytics and numerical time-series data analysis (NTSDA), where raw multivariate KPI windows from 4G/5G cell sites expand into thousands of floating-point tokens. Vision-Language Models (VLMs) eliminate this mismatch by encoding time-series as 2D plots, achieving 3.6-10.4x input token reduction across Llama-3.2-90B, Qwen2.5-VL-72B, and Pixtral-12B architectures. This translates to 1.8-2.5x measured inference energy reduction, saving approximately 7.2 MJ/day at telecom edge deployments and CloudRAN that monitor 200 cells per 15-minute interval. Critically, efficiency gains do not sacrifice accuracy: a fine-tuned Llama-3.2-90B-Vision VLM achieves 220.7% higher precision than its text-only counterpart and outperforms LSTM and ARIMA baselines by over 144% on telecom anomaly detection. On public benchmarks, Pixtral-12B achieves a 20.6x improvement in J/F1 score at mean F1 = 0.82. At 24 KPIs, text representations exceed the 128K context window of most production LLMs, rendering text-only processing infeasible without truncation, while visual representations remain within standard limits. These results establish VLMs as an energy-efficient and accuracy-superior modality for numerical time-series workloads, providing empirical grounding for AI inference systems that treat energy consumption as a first-class engineering constraint.