arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

代理验证的大语言用户体验微模拟:一种用于早期决策支持的工件优先协议

Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Decision Support

Alexandre Cristovão Maiorano

arXiv 2608.13563首次发表:更新:

发表机构

Lumytics(卢米提克斯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出代理验证的LLM驱动UX微模拟流程,通过代理语料库验证模拟结果,结合多指标对比基线、消融实验分析智能体策略,提供可复现的早期UX决策支持方案。

AI 中文摘要

早期团队往往缺乏用户、时间和预算来开展多次用户体验(UX)研究,但仍需要面向决策的信号以安全迭代。本文研究一种由大语言模型(LLM)驱动的UX微模拟流程,该流程可从版本化提示词、用户画像、任务和UI快照生成结构化的客户体验反馈(包括演练步骤、摩擦点、微调查信号)。由于带有任务结果的公共可用性数据集稀缺,我们使用多个公共代理语料库(应用评论、支持推文和开源软件问题)验证模拟的摩擦主题。我们提出一种轻量的代理验证协议,包含两个对齐指标:top-k Jaccard和分布加权Jaccard(W),并在六个代理数据集上对比词汇、TF-IDF和多语言嵌入基线。基于嵌入的对齐在主要应用评论和支持推文代理上产生的W高于词汇基线(例如,在Gojek上W=0.128对比0.000),同时表明top-k Jaccard在k值较大时会高估对齐效果。我们在Azure OpenAI部署上对四种智能体策略(单轮、N选优、混合以及本文提出的评分后选择评判器)进行消融实验,并报告8组方法-数据集对的自助法置信区间;这些区间显示,在我们的子样本规模下,嵌入W的点估计在重采样时存在系统性不稳定。我们还对接地和生成代理进行失败模式分析,记录校准注意事项以及被对抗性评判器标记为生成的输出的示例。本文的工件优先流程可从版本化运行工件生成可复现的表格和图表,支持在最终付费模型校准前进行迭代提示词和分类法优化。

英文摘要

Early-stage teams often lack users, time, and budget to run repeated UX studies, yet still need decision-oriented signals to iterate safely. We study an LLM-driven UX micro-simulation pipeline that generates structured customer-experience feedback (walkthrough steps, friction points, micro-survey signals) from versioned prompts, personas, tasks, and UI snapshots. Because public usability datasets with task outcomes are scarce, we validate simulated friction themes using multiple public proxy corpora (app reviews, support tweets, and open-source software issues). We propose a lightweight proxy-validation protocol with two alignment metrics: top-k Jaccard and distributional weighted-Jaccard (W), and compare lexical, TF-IDF, and multilingual embedding baselines across six proxy datasets. Embedding-based alignment yields higher W than lexical baselines on primary app-review and support-tweet proxies (e.g., W=0.128 vs 0.000 on Gojek), while top-k Jaccard is shown to overstate alignment at large k. We ablate four agent strategies (single-pass, best-of-N, hybrid, and a proposed score-then-select judge) across Azure OpenAI deployments and report bootstrap confidence intervals over 8 method-dataset pairs; these intervals reveal that the embedding W point estimate is systematically unstable under resampling at our subsample size. We also provide a failure-mode analysis of grounding and fabrication proxies, with documented calibration caveats and worked examples of outputs flagged as fabricated by an adversarial judge. Our artifact-first pipeline produces reproducible tables and figures from versioned run artifacts, supporting iterative prompt and taxonomy refinement before final paid-model calibration.

Comments27 pages, 5 figures, 15 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑