arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10412cs.HCcs.CY

当面试官是机器人时:多模态大语言模型(MLLM)主导的面试中的行为、故障与信任

When the Interviewer Is a Bot: Behavior, Breakdowns, and Trust in MLLM-Led Interviews

He Zhang, Kambinachi Chukwuma, ChanMin Kim, John M. Carroll

AI总结:

该研究通过构建InterviewBot系统开展实证研究,分析了MLLM主导面试的行为、故障与社会动态,为以人为中心的面试自动化提供设计启示。

AI中文摘要:

半结构化访谈是定性研究的基石,但仍需耗费大量人力。我们报告一项实证研究,探究当面试官是现成的实时多模态大语言模型(MLLM)时实际会发生什么。我们构建了InterviewBot,这是一个基于语音的面试系统,它将实时MLLM与研究者编写的大纲封装在一起,并非作为新架构,而是作为研究工具,用于观察默认MLLM的面试行为。在一项实践研究中(N=15),参与者完成了由机器人主导的半结构化访谈,随后进行了由人类主导的关于该体验的反思环节。我们贡献了:(i)对MLLM面试官的逐轮行为分析(N_轮次=428),显示其以回应为主,但深度提问不足(深化提问占所有轮次的4.9%),且28.7%的含问题轮次在一轮中包含多个问题,尽管有明确的“一次一个问题”的指令;(ii)在已部署而非模拟的系统中观察到的四类数据收集故障的归纳目录:信息丢失、提前终止、延迟和中断;(iii)参与者反思中呈现的三种社会动态:披露校准,即社会压力降低与阐述程度较浅同时出现;制度合法性,即信任与感知到的风险以及向AI授权所传递的组织者的信息相关,而非对话能力;对话接地,即内容相关的复述,而非通用的社交填充语,是参与者所认为的“倾听”。我们最后得出以人为中心的面试自动化在深度控制、透明交接和非模板化倾听机制方面的设计启示。

英文摘要:

Semi-structured interviews are a cornerstone of qualitative research but remain labor-intensive. We report an empirical study of what actually happens when the interviewer is an off-the-shelf real-time multimodal LLM (MLLM). We built InterviewBot, a voice-based interviewing system that wraps a real-time MLLM with a researcher-authored outline, and deployed it not as a novel architecture but as a research instrument for observing default MLLM interviewing behavior. In a practice study (N=15), participants completed a bot-led semi-structured interview and then a human-led reflection session about that experience. We contribute (i) a turn-level behavioral analysis of an MLLM interviewer (N_turns=428) showing that it is acknowledgment-heavy but probe-light (deepening probes account for 4.9% of all turns), and that 28.7% of question-bearing turns pack multiple questions into one turn despite an explicit one-question-at-a-time instruction; (ii) an inductive catalogue of four data-collection breakdowns (information loss, premature termination, latency, and interruption) observed in a deployed rather than simulated system; and (iii) three social dynamics from participants' reflections: disclosure calibration, where reduced social pressure coincided with shallower elaboration; institutional legitimacy, where trust tracked perceived stakes and what delegation to AI signaled about the organizer rather than conversational competence; and conversational grounding, where content-grounded paraphrase, not generic social filler, was what participants read as listening. We conclude with design implications for depth control, transparent handoffs, and non-templated listening mechanisms in human-centered interview automation.

补充信息

↑