arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

E-AVI:面向自动化视频面试的基于证据的多模态评估

E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews

Haoshen Wang, Dongbo Che, Zeyi Xie, Yuanjie Du, Shicheng Hua, Xingyu Wang

arXiv 2609.20001首次发表:更新:

发表机构

The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

E-AVI框架通过提取带时间戳的多模态证据并结合维度条件注意力,提升视频面试评分的预测性能,同时提供可检查的反馈与问答支持。

AI 中文摘要

自动化视频面试评估整合了语言内容、声学表达和视觉行为,然而仅靠数值预测所提供的可检查支持有限。我们提出了E-AVI,一个基于证据的框架,该框架提取带时间戳的多模态证据,并将维度条件证据注意力与源级嵌入相结合以进行评分。一个共享的证据池进一步支持自然语言反馈和后续问题回答。在RecruitView和一个私有酒店数据集上,E-AVI在秩相关方面始终优于微调的多模态基线。消融、证据删除、自助法、人工审计和问答分析表征了证据路径的预测贡献、依据性和实际效用。总之,这些结果表明,我们提出的E-AVI框架在提高预测性能的同时,为评估、反馈和交互式分析提供了可检查的支持。

英文摘要

Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑