发表机构
The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
E-AVI框架通过提取带时间戳的多模态证据并结合维度条件注意力,提升视频面试评分的预测性能,同时提供可检查的反馈与问答支持。
AI 中文摘要
自动化视频面试评估整合了语言内容、声学表达和视觉行为,然而仅靠数值预测所提供的可检查支持有限。我们提出了E-AVI,一个基于证据的框架,该框架提取带时间戳的多模态证据,并将维度条件证据注意力与源级嵌入相结合以进行评分。一个共享的证据池进一步支持自然语言反馈和后续问题回答。在RecruitView和一个私有酒店数据集上,E-AVI在秩相关方面始终优于微调的多模态基线。消融、证据删除、自助法、人工审计和问答分析表征了证据路径的预测贡献、依据性和实际效用。总之,这些结果表明,我们提出的E-AVI框架在提高预测性能的同时,为评估、反馈和交互式分析提供了可检查的支持。
英文摘要
Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.