arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM4OSC:用于开放声音控制的具有确定性验证的配置文件约束自然语言控制

LLM4OSC: Profile-Bound Natural Language Control with Deterministic Validation for Open Sound Control

Yuan-Yi Fan

arXiv 2607.26024首次发表:更新:

AI 中文总结

研究针对开放声音控制中语言模型的问题,提出LLM4OSC本地优先架构,通过人工审核配置文件等方式进行验证等操作,在特定配置文件测试中后端表现良好,主张将提议-验证-发送和错误发送率作为关键指标。

AI 中文摘要

开放声音控制(OSC)是专业音频、现场表演和虚拟制作中实时参数控制的主要有线协议。大语言模型可以生成看似合理的OSC,但会产生错误的地址、处理不当类型标签,在关键场景下无法使用。我们提出了LLM4OSC,这是一种本地优先架构,模型通过人工审核的设备配置文件提出结构化意图JSON,确定性代码在任何UDP发送之前进行验证、钳位和编码。我们引入了一个带有错误发送率CI门的冻结评估工具。在Max/MSP英雄配置文件上,经过配置文件标签丰富、符号插槽填充、自然语言细化和检索置信度门后,后端B0--B3都通过了冻结门。我们主张将提议-验证-发送和错误发送率作为语言到控制系统的一流指标。

英文摘要

Open Sound Control (OSC) is the dominant wire protocol for real-time parametric control in professional audio, live performance, and virtual production. Large language models can emit plausible OSC, but they hallucinate addresses, mishandle type tags, and fail under paraphrase- unacceptable in show-critical contexts. We present LLM4OSC, a local-first architecture in which models propose structured intent JSON over a human-reviewed device profile, and deterministic code validates, clamps, and encodes before any UDP send. We introduce a frozen evaluation harness with CI gates on wrong-send rate: mismatches that would still pass validation and transmit. On a Max/MSP hero profile (12 patterns; 8 literal + 8 paraphrase + 4 refusal cases), after profile tag enrichment, symbolic slot fill, NL refine, and a retrieval confidence gate, backends B0--B3 all pass frozen gates (100% semantic accuracy, 0% wrong-send). B0 (rules) remains the production default at ~0.05ms; LLM backends remain ~3-4s. Historical few-shot B2 accuracy of 62.5% rises to 100% on this suite only after symbolic post-processing- not because the 0.5B model alone becomes show-safe. We argue for propose-validate-send and wrong-send rate as first-class metrics for language-to-control systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑