arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35833cs.AIcs.CL

神经符号路由:在资源受限边缘设备上实现可靠推理

Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices

Avyay Sadhu, Alvaro Velasquez, Lekai Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对边缘设备上小型语言模型在结构化任务中不可靠的问题,提出神经符号路由器,利用L*算法学习确定性有限自动机,将查询分派给符号求解器或SLM,在树莓派上实现98.3%的准确率并显著提升速度与能效。

中文摘要 AI 辅助

在边缘硬件上运行语言模型可以在无网络连接的情况下提供私密且低延迟的推理,然而,适合此类设备的小型模型在计算机应擅长处理的任务(如算术、代数和形式逻辑问题)上并不可靠。我们认为,这种不可靠性在很大程度上是可以避免的。许多看似需要推理的查询实际上在结构上是确定性的,并且允许快速且精确的符号求解。因此,强迫概率模型去近似这些查询,牺牲了准确性和能量却收益甚微。我们提出了一种神经符号路由器,它对每个传入查询进行分类,并将其分派给最便宜的正确答案求解器,将结构化任务发送给确定性引擎,并将小型语言模型(SLM)保留用于开放式文字问题。我们不是手工编码路由逻辑,而是使用L*语法推断算法学习一个确定性有限自动机(DFA),以SLM作为成员资格预言机,以标记数据作为等价预言机。在树莓派4B(8 GB RAM,无GPU)上,针对来自DeepMind Mathematics、GSM8K和RuleTaker的100个未测试提示进行评估,学习到的路由实现了100%的路由准确率和98.3%的总体准确率,推理预算为512个令牌(文字问题准确率为93.3%),而最强智能体基线Program-of-Thought的准确率为72.0%,使用相同求解器的工具调用智能体的准确率为58.7%。由于格式化查询永远不会到达模型,路由器在1-11毫秒内回答这些查询,并且在其30令牌配置下,运行速度比Program-of-Thought快8.8倍,能效高2.8倍。

英文摘要

Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra, and formal logic problems. We argue that much of this unreliability is avoidable. Many queries appearing to demand reasoning are in fact structurally deterministic and permit fast and exact symbolic solutions. Therefore, forcing a probabilistic model to approximate them sacrifices accuracy and energy for little benefit. We present a neurosymbolic router that classifies each incoming query and dispatches it to the cheapest correct solver, sending structured tasks to deterministic engines and reserving the small language model (SLM) for open-ended word problems. Instead of hand-coding the routing logic, we learn a deterministic finite automaton (DFA) with the L* grammatical inference algorithm, using the SLM as a membership oracle and labeled data as an equivalence oracle. On a Raspberry Pi 4B (8 GB RAM, no GPU), evaluated on 100 untested prompts from DeepMind Mathematics, GSM8K, and RuleTaker, learned routing attains 100% routing accuracy and 98.3% overall accuracy with a 512-token reasoning budget (93.3% on word problems), compared with 72.0% for the strongest agent baseline, Program-of-Thought, and 58.7% for a tool-calling agent given the same solvers. Since formatted queries never reach the model, the router answers them in 1-11 ms and, in its 30-token configuration, runs 8.8x faster and 2.8x more energy-efficient than Program-of-Thought.

发表机构

  • University of Colorado Boulder(科罗拉多大学博尔德分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑