arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SERUM:面向用户建模的状态提取与优化

SERUM: State Extraction and Refinement for User Modeling

Andy J. Phu, Karin de Langis, James Mooney, Khanh Chi Le, Dongyeop Kang

arXiv 2607.29181首次发表:更新:

发表机构

University of Minnesota(明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SERUM是首个无需人工标注即可从非结构化自我中心屏幕视频生成可解释过程模型的多阶段框架,经61段跨领域视频验证,其优化后的标签在马尔可夫模型预测与人工评估中均表现优异,为用户建模提供了可扩展方案。

AI 中文摘要

具备主动个性化交互能力的智能助手需要结构化的用户意图与工作流模型,但从原始非结构化屏幕活动中构建这类模型仍是未解决的挑战。我们提出SERUM,一种多阶段框架,可利用分层视觉语言模型(VLM)标注直接从非结构化自我中心视频中提取有限状态行为模型。通过滑动窗口处理屏幕录制内容,SERUM在活动识别与意图推理阶段交替执行,每个阶段利用累积的先验上下文优化标签,以减少单阶段标注中出现的幻觉与时间混淆问题。随后,通过句子嵌入与人工校准阈值将同义状态合并为紧凑连贯的分类体系。我们通过在生成的标签序列(含动作与意图)上拟合一阶马尔可夫模型,并测量其相对于频率基线的预测准确率来评估行为结构。在四个领域(编码、烹饪、体育活动、日常生活)的61段自我中心视频上,我们发现:(1)迭代标签优化在数轮后收敛至稳定状态词汇,我们将其称为示意图平衡;(2)归一化马尔可夫模型的困惑度显著低于频率基线,动作预测准确率更高,在编码等结构化任务上增益最大;(3)人工标注者认为最终阶段标签准确,且相比第一阶段标签有显著改进。据我们所知,SERUM是首个无需人工标注即可从非结构化自我中心屏幕视频生成可解释过程模型的系统,为野外场景下的用户建模与行为理解开辟了可扩展路径。我们的演示、代码与结果均公开可用。

英文摘要

Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. However, building these models from raw, unstructured screen activity remains an open challenge. We present SERUM, a multi-pass framework that extracts finite-state behavioral models directly from unstructured egocentric video using hierarchical VLM annotation. Processing screen recordings through a sliding window, SERUM alternates between activity-recognition and intent-inference passes, with each pass refining labels using accumulated prior context to reduce hallucination and temporal conflation seen in single-pass annotation. Synonymous states are then merged via sentence embeddings and human-calibrated thresholds into a compact, coherent taxonomy. We evaluate behavioral structure by fitting first-order Markov models over the resulting label sequences (both actions and intents) and measuring predictive accuracy against frequency baselines. Across 61 egocentric videos in four domains (coding, cooking, physical activities, and daily life), we find: (1) iterative label refinement converges to a stable state vocabulary, which we term schematic equilibrium, after several passes; (2) normalized Markov models achieve substantially lower perplexity and higher action predictions than frequency baselines, with the largest gains on structured tasks like coding; and (3) human annotators rate final-pass labels as accurate and meaningfully improved over first-pass labels. To our knowledge, SERUM is the first system to produce interpretable process models from unstructured egocentric screen video without manual annotation, opening a scalable pathway for user modeling and behavioral understanding in the wild. Our demo, code, and results are publicly available

CommentsPublished as a conference paper at COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑