arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

工作节奏:面向职场智能体的人类行为轨迹多尺度解读

Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents

Lin Ai, Scott Counts

arXiv 2609.04556首次发表:更新:

发表机构

Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对职场智能体,提出构建多分辨率人类行为轨迹解读框架,在生产力套件数据上验证其结构稳定性与预测有效性,表明需按问题需求选择时间粒度解读轨迹。

AI 中文摘要

运行时轨迹正成为理解智能体系统的核心载体,但现有解读大多聚焦于智能体的行为。职场智能体面临的互补问题是:解读围绕它们的人类活动。数小时的低层级事件承载着关于用户状态的丰富证据,但粒度过细难以直接推理;将其展平为单一序列或压缩为单个嵌入,都把“总结用户行为”视为只有一个正确答案。我们则认为,行为解读依赖于分辨率:同一轨迹应能在不同时间分辨率下提供多个可被处理的解读。我们构建了一个多分辨率词汇表,包含语义归一化算子、重复 motif、连贯片段和日级节奏,每类都保留了自身时间尺度下的显著结构。将其应用于来自某大型商业生产力套件的6.67亿条人类归因事件(涉及5万名用户、100个组织),共得到120种算子类型、数千个motif、25种片段类型和5种日节奏原型。我们通过真实遥测数据验证:在2000名用户的不相交样本上重新运行整个流程,能恢复相同的分类体系(结构稳定性);在留存用户上,完整表示预测用户下一个片段的准确性优于单一算子基线,相对宏F1提升17%(预测有效性),说明这些抽象不仅能描述,还保留了与未来相关的信息。控制分辨率的消融实验显示,没有单一分辨率能适配所有问题:同一轨迹的不同智能体相关问题,在不同分辨率下能得到最佳解答。因此,面向智能体的行为轨迹解读应采用多分辨率且查询条件化的方式:智能体应能访问问题所需的时间粒度,而非采用单一通用总结。

英文摘要

Runtime traces are becoming a central substrate for understanding agentic systems, yet interpretation has focused largely on what the agent did. Workplace agents face the complementary problem: interpreting the human activity that surrounds them. Hours of low-level events carry rich evidence about a user's state but are too granular to reason over directly, and flattening them into one stream or compressing them into a single embedding both treat "summarize the user's behavior" as if it had one correct answer. We argue instead that behavioral interpretation is resolution-dependent: the same trace should admit multiple addressable interpretations at different temporal resolutions. We construct a multi-resolution vocabulary of semantically normalized operators, recurring motifs, coherent episodes, and day-level rhythms, each preserving the structure salient at its own horizon. Applied to 667 million human-attributed events from a large commercial productivity suite (50,000 users, 100 organizations), it yields 120 operator types, thousands of motifs, 25 episode types, and five day-rhythm archetypes. We validate it on real telemetry: re-running the entire pipeline on a disjoint 2,000-user sample recovers the same taxonomy (structural stability), and on held-out users the full representation forecasts a user's next episode more accurately than a flat-operator baseline, a 17% relative macro-F1 gain (predictive validity), so the abstractions preserve future-relevant information rather than merely describe it. A controlled resolution ablation then shows that no single level is optimal across questions: different agent-facing questions about the same trace are best answered at different resolutions. Behavioral trace interpretation for agents should therefore be multi-resolution and query-conditioned: an agent should access the temporal grain a question needs, not one universal summary.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑