实证软件工程中的人工智能主导访谈:一份经验报告
AI-Conducted Interviews in Empirical Software Engineering: An Experience Report
AI总结:
该经验报告探讨定制MyGPT在实证软件工程研究中主导访谈的应用,介绍其流程,分析提交工件和问卷回复,发现多数参与者评价积极,但存在局限,表明人工智能访谈可作补充选项,需相关设计与监督。
AI中文摘要:
半结构化访谈在实证软件工程中广泛使用,但资源密集且跨日程、地点和自然语言协调困难。本经验报告考察了定制的MyGPT在两项实证软件工程研究中进行简短的自我管理访谈的情况,一项关于重构实践,另一项关于Scrum相关活动中的生成式人工智能。参与者通过共享链接访问访谈者,使用语音交互,选择偏好的自然语言,在无研究人员在场的情况下完成访谈。人工智能遵循预定义协议生成结构化综合内容,参与者自愿提交。分析了66份提交内容和问卷回复,审核了工件格式、语言、长度和协议一致性。结果显示大部分提交工件符合格式,语言多为葡萄牙语,参与者对体验评价积极,但也存在一些局限。研究结果支持了该工作流程在受访者中的可行性和可接受性,但未确定完成率、时间节省、总结保真度或与人工访谈的等效性。因此,人工智能访谈者应被视为短期、聚焦、低风险研究的补充选项,并需进行协议设计、隐私指导、工件验证和人工监督。
英文摘要:
Semi-structured interviews are widely used in empirical software engineering (ESE), but they are resource-intensive and difficult to coordinate across schedules, locations, and natural languages. This experience report examines a customized MyGPT used to conduct short, self-administered interviews in two ESE studies: one on refactoring practices and another on generative AI in Scrum-related activities. Participants accessed the interviewer through shared links, used voice interaction, selected a preferred natural language, and completed the interview without a researcher present. The AI followed a predefined protocol and generated a structured synthesis that participants voluntarily submitted; these artifacts were not treated as verbatim transcripts. We analyzed 66 submissions and questionnaire responses, and audited artifact format, language, length, and protocol consistency. Of the submitted artifacts, 92.4% followed the expected synthesis format, 65 were predominantly in Portuguese and one in English, and two conflicted with the reported protocol. Participants generally rated the experience positively: 90.9% reported a positive overall experience and comfort, 95.5% considered the questions clear, 97.0% rated the pace positively, and 89.4% would participate again. Reported limitations included generic questions, limited sensitivity to answers, insufficient depth, privacy concerns, and missed human interaction. The findings support the operational viability and acceptability of this workflow among analyzed respondents, but do not establish completion rates, time savings, summary fidelity, or equivalence to human-conducted interviews. AI interviewers should therefore be treated as a complementary option for short, focused, low-risk studies, with protocol design, privacy guidance, artifact validation, and human oversight.