AI 中文总结
OmniEye是一个多模态视频智能系统,利用基础模型联合感知执法随身摄像头视频和音频,通过SQLite与BM25支持智能体查询,在16GB GPU上高效运行,支持执法培训与审查。
AI 中文摘要
我们介绍了OmniEye,一个用于执法培训和审查的多模态视频智能系统(源代码可应请求提供给经过验证的执法和公共安全机构)。OmniEye摄入随身摄像头拍摄的镜头,并通过一个多模态基础模型联合感知每个30秒窗口的视频和音频。然后,它将模型的结构化输出存储在带有BM25全文搜索的嵌入式SQLite数据库中。警官可以通过一个智能体来查询镜头,该智能体编写结构化查询,检索候选窗口,并在可能引用它们之前用模型重新感知它们。整个系统运行在一块16 GB GPU上,采用4位量化感知训练模型,并且还可以在多GPU集群上扩展到完整的bf16精度。
英文摘要
We introduce OmniEye, a multimodal video intelligence system for law-enforcement training and review (source code available on request to verified law-enforcement and public-safety agencies). OmniEye ingests body-worn camera footage and perceives every 30-second window jointly across video and audio with one multimodal foundation model. It then stores the model's structured output in an embedded SQLite database with BM25 full-text search. Officers can question the footage through an agent that writes structured queries, retrieves candidate windows, and re-perceives them with the model before it may cite them. The whole system runs on one 16 GB GPU with a 4-bit quantization-aware-trained model, and it also scales to full bf16 precision on a multi-GPU cluster.