面向智能手术室的语音交互多智能体系统:架构设计与关键技术
A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies
浏览论文内容
中文总结 AI 辅助
本文提出SurgicalRoomAgent,一种基于大语言模型的智能手术室语音交互多智能体系统,通过分层架构实现自然语言理解与设备控制,并采用KV缓存前缀预热、流式部分JSON解析和渐进式技能提示三项关键技术降低延迟、提升效率,实验验证满足实时性要求。
中文摘要 AI 辅助
本文介绍了SurgicalRoomAgent,一个基于大语言模型(LLMs)的智能手术室语音交互多智能体系统。该系统通过分层架构实现自然语言理解、设备控制、术中记录和手术报告生成,该架构包括语音交互流水线(唤醒、ASR、话轮检测、智能体推理、TTS)和智能体核心(技能注册表、任务规划器、设备管理器)。研究了三项关键技术:(1)用于低延迟推理的KV缓存前缀预热,通过字节级最长公共前缀复用,将重计算开销从约500毫秒降至几十毫秒;(2)流式部分JSON解析与早期并行任务执行,将端到端延迟降低约30%;(3)渐进式技能提示披露,根据用户角色、连接设备和手术阶段动态过滤系统提示,以在有限的上下文窗口内最大化信息密度。该系统使用Qwen3-27B模型及此http URL推理引擎实现。实验分析表明,系统在16,384个令牌的上下文限制内有效运行,多设备并行控制响应时间满足手术室实时性要求。
英文摘要
This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs). The system achieves natural language understanding, device control, intraoperative recording, and surgical report generation through a layered architecture comprising a voice interaction pipeline (wake, ASR, turn detection, agent reasoning, TTS) and an agent core (skill registry, task planner, device manager). Three key technologies are investigated: (1) KV Cache prefix warming for low-latency inference, reducing recomputation overhead from approximately 500 ms to tens of milliseconds via byte-level Longest Common Prefix reuse; (2) streaming partial JSON parsing with early parallel task execution, reducing end-to-end latency by approximately 30%; and (3) progressive skill prompt disclosure, which dynamically filters system prompts based on user role, connected devices, and surgical phase to maximize information density within limited context windows. The system is implemented using the Qwen3-27B model with llama.cpp/sglang inference engines. Experimental analysis demonstrates effective operation within a 16,384-token context limit and multi-device parallel control response times meeting OR real-time requirements.
发表机构
- Wuhan United Imaging Surgical Co., Ltd. (UIS)(武汉联影外科医疗科技有限公司(UIS))
机构由 AI 辅助整理,请以论文原文为准。