arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34180cs.AIcs.CV

面向文本介导视频异常检测的决策读出:Jev与Qwen的探索性评估

Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwen

Xukui Qin, Youting Wang, Xinjie He, Ziyang Luo, Runxiong Wu, Yan-Syuan Chen, Zhongyao Chu

首次发表
浏览论文内容

中文总结 AI 辅助

本文探索性评估了文本介导视频异常检测中Jev与Qwen决策读出的影响,发现Jev在XD字幕上平均精度显著更高,但UCF上无优势,且未证实因果性优势或普遍优越性。

中文摘要 AI 辅助

当视频衍生的文本证据保持不变时,决策读出(decision readout)的影响有多大?我们在来自UCF-Crime和XD-Violence的稀疏开发样本(包含40个视频和400个目标锚点)上评估了Jev类型化决策和三种Qwen读出,每个样本均以摘要和有序字幕形式呈现。每个数据集贡献20个源组和200个锚点,其中分别仅包含10个和37个正样本。最初的五后端试点要求4000次预测;Jev Choice在研究的严格数值策略下,从800次中返回了776个有效响应,从而阻碍了其全覆盖质量比较。在XD字幕上,Jev Noul实现了75.99%的平均精度,而Qwen生成概率为48.47%,更强的局部序数似然期望为57.81%。后者的配对差异为18.18个百分点(95%源组自助法区间5.53-31.50)。UCF未显示相应优势:字幕ROC-AUC在Noul上为52.26%,在序数似然上为65.95%。两种概率读出的UCF Brier分数均高于评估流行率参考值0.0475,因此更差。我们还对历史LAVAD分数在完全匹配的锚点上进行了审计,并将响应结构与数值一致性区分开来。缺少二元似然对照。这些探索性离线结果刻画了排序、概率质量和接口失败;它们既未确立因果类型化接口优势,也未确立普遍优越性、校准性或端到端加速。

英文摘要

How much does the decision readout matter when video-derived textual evidence is held fixed? We evaluate Jev typed decisions and three Qwen readouts on a sparse development sample of 40 videos and 400 target anchors from UCF-Crime and XD-Violence, each presented as a summary and ordered captions. Each dataset contributes 20 source groups and 200 anchors, including only 10 and 37 positives, respectively. The original five-backend pilot requested 4,000 predictions; Jev Choice returned 776 valid responses out of 800 under the study's strict numerical policy, blocking its full-coverage quality comparison. On XD captions, Jev Noul achieved 75.99% average precision versus 48.47% for Qwen generated probability and 57.81% for the stronger local ordinal-likelihood expectation. The latter paired difference was 18.18 percentage points (95% source-group bootstrap interval 5.53-31.50). UCF did not show a corresponding advantage: caption ROC-AUC was 52.26% for Noul and 65.95% for ordinal likelihood. Both probability readouts had higher, hence worse, UCF Brier scores than the evaluation-prevalence reference of 0.0475. We additionally audit historical LAVAD scores at exactly matched anchors and distinguish response structure from numerical consistency. A binary-likelihood control is missing. These exploratory offline results characterize ranking, probability quality and interface failures; they establish neither a causal typed-interface benefit nor general superiority, calibration or end-to-end acceleration.

补充信息

↑