arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AnovaX:一个具有大语言模型规划、类型化执行器和自适应恢复功能的本地多智能体语音助手

AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery

Raunak B Sinha

arXiv 2607.15367首次发表:更新:

发表机构

BITS Pilani(贝拉理工学院皮拉尼分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究桌面语音助手局限,提出AnovaX本地多智能体语音助手,通过Python进程连接多组件,含规划器、安全层等,各工具对应专用智能体类,有递归元智能体及自适应恢复循环,还介绍配套服务器功能,展示其独特优势。

AI 中文摘要

桌面语音助手仍由将原始音频传输到机器外部并提供固定技能集的云管道主导。我们介绍了AnovaX,一个小型的本地优先助手,它完全在用户计算机上运行,并将桌面本身作为其操作界面。一个Python进程将唤醒词门、语音管道、发出工具调用JSON计划的大语言模型规划器(Gemini)、白名单和黑名单安全层、在有界线程池上将每个计划转换为类型化子智能体的多智能体协调器以及在核心步骤失败时接管的自适应恢复循环连接在一起。每个工具对应一个具有自己的超时、重试策略和共享资源锁的专用智能体类。一个递归元智能体允许规划器将子目标委托回自身,嵌套限制为两级。恢复循环使用紧凑的ReAct风格提示,并通过只读工具的推测执行隐藏Gemini的延迟。一个配套的Flask服务器通过本地WiFi公开一个便于手机访问的远程接口,实时将每个智能体生命周期事件镜像到手机,并通过MJPEG流回笔记本电脑屏幕,以便用户可以观看远程命令的执行情况。该项目的意义不在于与Siri或Alexa竞争,而在于表明一个清晰的、几千行的助手足以打开应用程序、在其中输入、运行搜索、协调并发操作、从单步失败中恢复,并完全由另一个房间的手机驱动——而大语言模型无需接触键盘。

英文摘要

Desktop voice assistants are still dominated by cloud pipelines that ship raw audio off the machine and expose a fixed set of skills. We describe AnovaX, a small local-first assistant that runs entirely on the user's computer and treats the desktop itself as its action surface. A single Python process wires together a wake-word gate, a speech pipeline, an LLM planner (Gemini) that emits a JSON plan of tool calls, a whitelist-and-denylist safety layer, a multi-agent orchestrator that translates each plan into typed child agents on a bounded thread pool, and an adaptive recovery loop that takes over whenever a core step fails. Every tool corresponds to a specialized agent class (AppAgent, TypingAgent, BrowserAgent and six others) with its own timeout, retry policy, and shared-resource locks. A recursive MetaAgent lets the planner delegate a sub-goal back to itself, capped at two levels of nesting. The recovery loop uses a compact ReAct-style prompt and hides Gemini's latency behind speculative execution of read-only tools. A companion Flask server exposes a phone-friendly remote over the local WiFi, mirrors every agent lifecycle event to the phone in real time, and streams the laptop's screen back over MJPEG so the user can watch remote commands land as they run. The point of the project is less to compete with Siri or Alexa than to show that a legible, few-thousand-line assistant is enough to open apps, type into them, run searches, coordinate concurrent actions, recover from single-step failures, and be driven entirely from a phone in another room -- without the LLM ever touching the keyboard.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑