arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语义动作图:用于智能体基础定位与体育精彩片段人类解读的共享表示

Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights

Tica Lin, Deepak Chandran, Gauri Jagatap, Chen Chen, Andrea Fanelli, David Gunawan, Josh Kimball

arXiv 2609.20768首次发表:更新:

发表机构

Dolby Laboratories; XPENG(杜比实验室; 小鹏汽车)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出语义动作图,一种轻量级领域模式,通过结构化表示体育比赛,同时支持智能体生成叙述性精彩片段和人类可视化查询解读,并在SportSAGE中验证其有效性。

AI 中文摘要

生成式智能体越来越多地被用于选择和叙述视频精彩片段,但它们通常基于非结构化或帧级表示进行操作。因此,其输出难以让观众验证并引导至个人偏好。我们提出了语义动作图,一种轻量级领域模式,将体育比赛表示为表演者、动作、接受者、时刻和状态节点,这些节点通过角色、时间和结果边连接。该模式展示了三个关键特性:1)连接的事件序列,2)共享的封闭词汇表,以及3)可寻址帧的时刻,使其能够同时服务于两个消费者:一个用于生成叙述性精彩片段的智能体流水线,以及一个供观众查询和检查相同结构的可视化界面。我们在SportSAGE中实例化了该模式,这是一个将四模块精彩片段流水线与图形界面配对的设计探针,并报告了来自12名足球迷的反馈。参与者对生成的精彩片段和叙述的质量感到满意,并使用图形界面来搜索、导航和解读比赛精彩片段。这些结果提供了早期证据,表明一个小的、人类可读的模式可以同时为智能体生成提供基础并支持人类解读。

英文摘要

Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, and state nodes connected by role, temporal, and outcome edges. The schema demonstrates three key properties: 1) connected event sequences, 2) a shared, closed vocabulary, and 3) frame-addressable moments, making it suitable to serve two consumers at once: an agentic pipeline that composes narrated highlights, and a visual interface through which viewers query and inspect the same structure. We instantiate it in SportSAGE, a design probe pairing a four-module highlight pipeline with a graph interface, and report feedback from 12 soccer fans. Participants were satisfied with the quality of the generated highlights and narratives, and used the graph interface to search, navigate, and interpret the match highlights. These results provide early evidence that one small, human-readable schema can ground agent generation and support human interpretation at the same time.

Comments5 pages, 3 figures, Accepted for publication at IEEE VIS 2026 Workshop on GenAI, Agents, and the Future of VIS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑