arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LadderTeam:双智能体阶梯式引出框架

LadderTeam: Dual-Agent Laddering Elicitation Framework

Manjushree Aithal, Alexander Kotz, James Mitchell

arXiv 2608.17029首次发表:更新:

发表机构

University of Colorado Anschutz(科罗拉多大学安舒茨校区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LadderTeam是基于双智能体LLM架构的自动化UX线框图访谈框架,采用三种探测策略,经216次模拟访谈验证,实现高收敛率与可执行响应匹配率,且无主题偏离。

AI 中文摘要

从终端用户处获取详细且可执行的软件需求,是软件产品或应用迭代开发的关键阶段。为确保收集到的反馈详细且可执行,软件团队可采用阶梯式访谈技术。该技术虽能保障软件反馈的粒度与可执行性,但存在若干局限:传统阶梯式访谈为人工流程,时间与经济成本高,可扩展性受限;访谈者需在深挖信息与应对受访者行为、文化约束间寻求平衡。为解决这些局限,本文提出LadderTeam——一个开放、可复现的框架,采用双智能体大语言模型(LLM)架构自动化UX线框图访谈。活跃的访谈者智能体可执行三种探测策略(ACV、5问法、JTBD),从可用性反馈评论中引出可执行软件需求;并行运行的背景评判智能体则评估探测-响应对,并触发实时护栏以防止主题偏离。为严格评估LLM阶梯式访谈且排除参与者差异的混淆影响,本文引入受控模拟方法论,利用脚本化真实转录本将探测质量隔离为唯一实验变量。在216次访谈中,LadderTeam实现99.1%的链收敛率、81.0%的真实可执行响应匹配率(其中不情愿性格受访者匹配率为86.1%,简洁性格受访者匹配率为75.9%),且所有运行均无主题偏离。所有评估代码、转录本、输入内容及实时演示平台将在论文接收后开源。

英文摘要

Eliciting detailed and actionable software requirements from end-users is a critical phase in the iterative development of a software product or application. To ensure the feedback collected is detailed and actionable, software teams can leverage the laddering interview technique. While effective for ensuring granular and actionable items from the software feedback, these interviews are subject to several limitations. They are traditionally a manual process associated with a time and financial burden, limiting scalability; interviewers must balance probing for depth while managing interviewee behavioral and cultural constraints. To address these limitations, we present \textbf{LadderTeam}, an open, reproducible framework that automates UX wireframe interviews using a dual-agent Large Language Model (LLM) architecture. An active interviewer agent executes one of three probing strategies (ACV, 5-Whys, and JTBD) to elicit actionable software requirements from usability feedback comments, while a concurrent background Judge agent evaluates probe-response pairs and triggers real-time guardrails to prevent topic drift. To rigorously evaluate LLM laddering without participant variance confounds, we introduce a controlled simulation methodology utilizing scripted ground-truth transcripts to isolate probe quality as the sole experimental variable. Across 216 interviews, \textbf{LadderTeam} achieved 99.1\% chain convergence and an 81.0\% ground-truth actionable response match (86.1\% reluctant personality, 75.9\% terse personality) with zero drift across all runs. All evaluation code, all transcripts, inputs, and a live demonstration platform will be open-sourced upon acceptance.

Comments4 pages, 1 figure, 2 tables, Accepted in ACM AI Summit 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑