InteractionBench:流式视频系统的实时交互基准
InteractionBench: A Real-Time Interaction Benchmark for Streaming Video Systems
- University of Pennsylvania(宾夕法尼亚大学)
- Stanford University(斯坦福大学)
- University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Princeton University(普林斯顿大学)
- Hong Kong Polytechnic University(香港理工大学)
- University of California, San Diego(加利福尼亚大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
InteractionBench 是一个针对流式视频系统的实时交互基准,评估模型、记忆和响应控制器在 812 个视频上的 1,060 次交互中的说话与沉默决策,发现现有系统在沉默合规性上存在显著缺陷,且离线评分无法预测在线行为。
AI中文摘要:
视频助手必须在其指令需要响应时说话,否则保持沉默。我们引入了一个基准,用于评估模型、记忆和响应控制器这一完整系统的这一决策。InteractionBench 涵盖了 812 个视频中的 1,060 次交互中的查询响应、事件触发和持续更新,包括 69 个负向流和 53 个将计数事件与相似误报配对的套件。它在视频时钟上对内容准确性、时间准确性和沉默合规性进行评分。及时的语音会牺牲各系统的沉默性。轮询的 Qwen3-VL-8B 达到 77.8 的时间准确性,但沉默合规性仅为 10.9。一个原生实时交互系统在 66.8 的时间准确性下达到 29.2 的沉默合规性,但在 89.9% 的负向流上发出响应。没有开源权重系统能通过三分之一的误报套件。更少的回复仅在选择性删除时有所帮助,因为随机删除只是用时间换取沉默。离线评分会遗漏这些失败,并错误预测在线行为。增加约束代价高昂,因为原生系统的控制器本身增加甚少,而代理系统仅在约 30 秒内增加它。页面:此 https URL 代码:此 https URL 数据:此 https URL
英文摘要:
A video assistant must speak when its instruction warrants a response and stay silent otherwise. We introduce a benchmark that evaluates this decision for the complete system of model, memory, and response controller. InteractionBench covers query responses, event triggers, and ongoing updates in 1,060 interactions over 812 videos, with 69 negative streams and 53 suites that pair counted events with look-alike near misses. It scores content accuracy, timing accuracy, and silence compliance on the video clock. Timely speech costs silence across systems. Polled Qwen3-VL-8B reaches 77.8 timing accuracy but 10.9 silence compliance. A native real-time interaction system reaches 29.2 silence compliance at 66.8 timing accuracy, yet emits on 89.9% of negative streams. No open-weight system clears a third of the near-miss suites. Fewer replies help only when chosen, as random deletion merely trades timing for silence. Offline scores miss these failures and mispredict online behavior. Adding restraint is costly, as the native system's controller adds little by itself and agentic systems add it only at about 30 s per poll.Project page: https://www.enxinsong.com/projects/interactionbench/ Code: https://github.com/Espere-1119-Song/InteractionBench Data: https://huggingface.co/datasets/InteractionBench/InteractionBench