发表机构
ALMA PHIL; Technical University of Munich; CNRS-AIST Joint Robotics Laboratory (JRL)(ALMA PHIL; 慕尼黑工业大学; 法国国家科学研究中心-日本产业技术综合研究所联合机器人实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对移动语音助手的弃权(不执行)问题,基于500多用户的真实数据构建VoxFallbacks数据集,对比评估不同模型,发现轻量嵌入分类器性能更优,为设计高效弃权机制提供经验。
AI 中文摘要
对用户输入的鲁棒理解是部署在真实环境中的语音助手的核心要求。在实际应用中,这些系统会遇到各类弃权(不执行)情况,原因包括音频输入嘈杂、转录错误、请求模糊、语句不完整或意外激活。现有系统通常会返回通用的弃权(不执行)消息,无法解决底层交互故障,还可能降低用户体验。我们针对日常环境中用于一般健康支持的已部署智能手表语音助手,研究其弃权(不执行)处理。分析基于500多名用户六个月的真实使用数据,得到包含3030条匿名、自然发生的触发弃权(不执行)语句的数据集。我们贡献了三点:(1)操作分类法及带标注的VoxFallbacks数据集;(2)实际部署约束下分类流程中不同模型的对比评估;(3)设计鲁棒且经济高效的弃权(不执行)机制的实用经验。结果显示,基于轻量嵌入的分类器在多数分类任务上优于更大的生成式模型,同时所需计算资源大幅减少。
英文摘要
Robust understanding of user input is a core requirement for voice assistants deployed in real-world environments. In practice, these systems encounter heterogeneous fallback situations caused by noisy audio input, transcription errors, ambiguous requests, incomplete utterances, or unintended activations. Existing systems typically respond with generic fallback messages, which do not resolve the underlying interaction failure and can degrade user experience. We study fallback handling in a deployed smartwatch-based voice assistant for general health support in everyday environments. Our analysis is based on six months of real-world usage data from more than 500 users, yielding a dataset of 3,030 anonymized, naturally occurring fallback-triggering utterances. We contribute (1) an operational taxonomy and the annotated VoxFallbacks dataset of these interactions, (2) a comparative evaluation of different models within a classification pipeline under practical deployment constraints, and (3) practical lessons for designing robust and cost-efficient fallback mechanisms. Results show that lightweight embedding-based classifiers outperform larger generative models on most classification tasks while requiring substantially fewer computational resources.
CommentsAccepted to EMNLP 2026