arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniEye:面向执法随身摄像头的多模态取证视频智能系统

OmniEye: Efficient Multimodal Forensic Video Intelligence for Law-Enforcement Body-Worn Cameras

Mamadou K. Keita, Angela Srbinovska, Anita Srbinovska, Nishka Desai, Isabella Zicari, P. Kwaku Sanaah-Faried, Sanjay Charitesh Makam, Wyatt Auten, Vivek Senthil, Hannah Desnick, Jonathan Bateman, Adrian Martin, Christopher Homan, John McCluskey, Ernest Fokoué

arXiv 2609.09460首次发表:更新:

AI 中文总结

OmniEye是一个多模态视频智能系统,利用基础模型联合感知执法随身摄像头视频和音频,通过SQLite与BM25支持智能体查询,在16GB GPU上高效运行,支持执法培训与审查。

AI 中文摘要

我们介绍了OmniEye,一个用于执法培训和审查的多模态视频智能系统(源代码可应请求提供给经过验证的执法和公共安全机构)。OmniEye摄入随身摄像头拍摄的镜头,并通过一个多模态基础模型联合感知每个30秒窗口的视频和音频。然后,它将模型的结构化输出存储在带有BM25全文搜索的嵌入式SQLite数据库中。警官可以通过一个智能体来查询镜头,该智能体编写结构化查询,检索候选窗口,并在可能引用它们之前用模型重新感知它们。整个系统运行在一块16 GB GPU上,采用4位量化感知训练模型,并且还可以在多GPU集群上扩展到完整的bf16精度。

英文摘要

We introduce OmniEye, a multimodal video intelligence system for law-enforcement training and review (source code available on request to verified law-enforcement and public-safety agencies). OmniEye ingests body-worn camera footage and perceives every 30-second window jointly across video and audio with one multimodal foundation model. It then stores the model's structured output in an embedded SQLite database with BM25 full-text search. Officers can question the footage through an agent that writes structured queries, retrieves candidate windows, and re-perceives them with the model before it may cite them. The whole system runs on one 16 GB GPU with a 4-bit quantization-aware-trained model, and it also scales to full bf16 precision on a multi-GPU cluster.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑