发表机构
Occidental College(西方学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将COPUS构建为多模态基准,提出基于MiniCPM-V-4.5和MLP的VISTA基线,在化学讲座测试中受限宏准确率达80.1%,同时发布相关工具与代码。
AI 中文摘要
视频-语言基准通常由数据集作者构建,且未发布可靠性统计数据,导致该构建的噪声底值未知。我们认为,多模态基准测试受益于那些已投入策略以确保可靠性的研究社区所采用的方法。我们以本科STEM课堂观察协议(COPUS)为例进行说明:这是一个拥有十年同行评审可靠性文献的24编码多标签观察工具。我们将COPUS重新构建为多模态基础模型的视频基准,它提供了一组密集的结构化标签(在50-90分钟的讲座中每2分钟一个24维二元向量)、经外部验证的词汇表,以及基于人类评估者的每个编码可靠性目标的成熟文献。我们评估语料库中的注释由5人组成的人类评估小组生成,其共识矩阵作为我们的参考。我们提出VISTA,这是一个在密集滑动窗口上运行MiniCPM-V-4.5的基线模型,用在冻结骨干网络顶部训练的轻量级多层感知器(MLP)头优化每个窗口的输出,并将得到的预测结果最大池化到2分钟的COPUS网格上。在三个保留的化学讲座中,VISTA达到了80.1%的受限宏准确率,而零样本变体为74.9%,最大的残余误差出现在视觉相似的教师编码和罕见的依赖音频的编码上。我们描述了三种系统失效模式(音频部分可观测性、细粒度小组作业区分、长尾召回),并在该httpsURL发布了基准工具、提示和基线代码。
英文摘要
Video-language benchmarks are usually constructed by the dataset authors without published reliability statistics, leaving the noise floor of the construct unknown. We argue that multimodal benchmarking benefits from methods taken from research communities that have already invested in strategies to ensure reliability. We illustrate the case with the Classroom Observation Protocol for Undergraduate STEM (COPUS): a 24-code multi-label observation instrument with a decade of peer-reviewed reliability literature. We recast COPUS as a video benchmark for multimodal foundation models, where it provides a dense set of structured labels (a 24-dimensional binary vector every 2 minutes across a 50-90 minute lecture), an externally validated vocabulary, and established literature that provides a per-code reliability target based on human evaluators. Annotations in our evaluation corpus are produced by a 5-person human-evaluator panel whose consensus matrix is our reference. We propose VISTA, a baseline that runs MiniCPM-V-4.5 over a dense sliding window, refines its per-window outputs with a lightweight multi-layer perceptron (MLP) head trained on top of the frozen backbone, and max-pools the resulting predictions onto the 2-minute COPUS grid. On three held-out chemistry lectures, VISTA reaches 80.1% restricted macro accuracy versus 74.9% for the zero-shot variant, with the largest residual errors on visually similar instructor codes and on rare audio-dependent codes. We characterize three systematic failure modes (audio-partial observability, fine-grained group-work discrimination, long-tail recall) and release the benchmark tooling, prompts and baseline code at https://github.com/ajfranck/VISTA.
CommentsDataMFM Workshop @ Computer Vision & Pattern Recognition (CVPR) 2026