arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

传感器-语言-动作模型

Sensor-Language-Action Models

Yuekai Xu, Zitao Shuai, Yuzhe Yang

arXiv 2610.08244首次发表:更新:

发表机构

University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该论文提出传感器-语言-动作(SLA)框架,以语言为接口统一传感器观测与动作,构建大规模基准并推出OpenSLA模型,在临床预测等任务上超越现有方法,支持零样本泛化。

AI 中文摘要

传感器不仅有助于理解世界,还有助于决定下一步该做什么。然而,现有的传感器模型大多止步于感知:它们识别状态或预测结果,而将动作建模为通过特定任务且通常封闭的标签空间来单独处理。我们引入了传感器-语言-动作(SLA)建模,这是一个将多模态传感器观测、自然语言和动作连接在统一模型中的框架。SLA使用语言作为感知与动作之间的语义接口,使得异构动作能够被表示、预测和解释,同时保持其基于底层传感器证据的根基。我们构建了一个大规模的SLA基准,包含跨越超过116,000名个体、79种传感器模态和60个动作组的数据集,以及一个多方面的描述流水线,该流水线对齐了用户上下文、传感器动态和动作证据。在此框架的基础上,我们提出了OpenSLA,一个用于层次化动作预测、状态理解和动作解释的统一SLA模型。在临床预测、手术室和代谢健康等真实世界任务上的大量实验验证了其相对于最先进方法的优越性能。OpenSLA还展示了引人入胜的能力,包括语言引导的证据基础和对未见动作及群体的零样本泛化。

英文摘要

Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models however largely stop at perception: they recognize states or predict outcomes, leaving actions modeled separately through task-specific and often closed label spaces. We introduce Sensor-Language-Action (SLA) modeling, a framework that connects multimodal sensor observations, natural language, and actions within a unified model. SLA uses language as a semantic interface between sensing and acting, allowing heterogeneous actions to be represented, predicted, and explained while remaining grounded in the underlying sensor evidence. We build a large-scale SLA benchmark consisting of datasets that span more than 116,000 individuals, 79 sensor modalities, and 60 action groups, together with a multi-faceted captioning pipeline that aligns user context, sensor dynamics, and action evidence. Building on this framework, we present OpenSLA, a unified SLA model for hierarchical action prediction, state understanding, and action explanation. Extensive experiments on real-world tasks in clinical prediction, operating rooms, and metabolic health verify its superior performance over the state-of-the-art. OpenSLA also demonstrates intriguing capabilities including language-guided evidence grounding and zero-shot generalization to unseen actions and cohorts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑