YazSes:一款离线、隐私优先的跨平台按键式语音听写系统
YazSes: An Offline, Privacy-First, Cross-Platform Hold-to-Talk Voice-Dictation System
浏览论文内容
中文总结 AI 辅助
YazSes是一款完全设备端运行的跨平台离线语音听写系统,采用faster-whisper实现高准确率转录,无数据外流,命令语法动作准确率100%,已开源且可复现。
中文摘要 AI 辅助
云语音听写服务准确率高,但需将用户语音流式传输至远程提供商,这在隐私敏感行业及离线、气隙环境中是不可接受的权衡;主流设备端替代方案要么受平台限制,要么面向专家脚本编写而非即插即用式听写。本文提出YazSes,这是一款开源(Apache-2.0许可)的按键式语音听写守护程序,完全在设备端运行,通过基于协议的平台抽象层实现单一代码库适配Linux、macOS和Windows。YazSes使用faster-whisper(CPU,int8精度)在本地转录语音,并将结果注入焦点应用;快速正则表达式命令语法(可选小型语言模型路由器支持)可将语音映射到编辑器和终端操作。无数据流出设备:录音采用按键式而非始终监听,无遥测数据,可选个性化循环将语料库加密存储在设备端,并提出配置更改而非传输数据。本文描述了系统架构——基于协议的平台抽象层后的分阶段流水线,带有JSON-RPC控制平面——及其隐私和威胁模型。我们在一台普通Linux笔记本电脑上评估了发布的Python实现;macOS和Windows后端已实现并通过单元测试,但此处未进行端到端评估。在涵盖40位说话者的200条LibriSpeech test-clean话语中,word error rate在2.59%至4.82%之间,实时因子为0.520,在无GPU的CPU上解码速度快于实时。命令语法在普通听写中达到100%动作准确率,假阳性率为0.0%,每次调用耗时0.021毫秒,非解码流水线增加0.289毫秒开销。本文所有数据背后的系统和可复现基准测试工具均已公开。
英文摘要
Cloud voice-dictation services deliver strong accuracy but require streaming a user's speech to a remote provider, an unacceptable trade-off in privacy-sensitive professions and offline or air-gapped settings; the leading on-device alternatives are either platform-locked or aimed at expert scripting rather than plug-and-play dictation. We present YazSes, an open-source (Apache-2.0) hold-to-talk voice dictation daemon that runs entirely on-device, with a single codebase targeting Linux, macOS, and Windows through a protocol-based platform abstraction. YazSes transcribes speech locally with faster-whisper (CPU, int8) and injects the result into the focused application; a fast regex command grammar, backed by an optional small-language-model router, maps utterances to editor and terminal actions. Nothing leaves the machine: recording is push-to-talk rather than always-listening, there is no telemetry, and an opt-in personalization loop keeps its corpus encrypted on-device and proposes configuration changes instead of shipping data out. We describe the system architecture -- a staged pipeline behind a protocol-based platform abstraction with a JSON-RPC control plane -- and its privacy and threat model. We evaluate the shipping Python implementation on a single commodity Linux laptop; the macOS and Windows backends are implemented and unit-tested but not end-to-end evaluated here. On 200 LibriSpeech test-clean utterances spanning 40 speakers, word error rate ranges from 4.82% (tiny.en) to 2.59% (small.en) at a real-time factor of 0.520 for small.en, decoding faster than real time on CPU with no GPU. The command grammar reaches 100% action accuracy with a 0.0% false-positive rate on plain dictation at 0.021 ms per call, and the non-decode pipeline adds 0.289 ms of overhead. The system and the reproducible benchmark harness behind every number in this paper are public.