arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用户如何看待生成式人工智能:应用商店评论中信任与摩擦的跨平台NLP分析

What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

Md Jafrin Hossain, Umme Nusrat Jahan, Shouvaggo Sharif Shammo

arXiv 2609.19151首次发表:更新:

发表机构

KFCIS, Florida International University; DoA, Bangladesh University of Engineering and Technology(佛罗里达国际大学; 孟加拉国工程技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分析六大GenAI应用的17,012条应用商店评论,结合主题建模与情感分类,揭示了广告、认证、服务器和订阅定价中的负面情绪集中,并发现Claude用户情绪显著极化,提出了信任摩擦评分以量化用户信任与可用性障碍。

AI 中文摘要

生成式人工智能(GenAI)应用已实现快速消费者采纳,然而少有大规模研究考察用户感知质量、信任和采纳障碍。我们提出了对六大主要GenAI应用(ChatGPT、Gemini、Microsoft Copilot、Claude、DeepSeek和Perplexity)的应用商店评论进行跨应用分析的首批研究之一,涵盖来自Google Play和Apple App Store的17,012条英文评论。我们结合BERTopic主题建模与RoBERTa情感分类,并使用卡方检验、Kruskal-Wallis检验和多项逻辑回归(经Bonferroni校正)评估跨应用差异。两个组件均基于300条评论的分层样本,通过人工编码进行验证。结果显示,负面情绪集中在广告(91%)、身份验证(89%)、服务器可靠性(83%)和订阅定价(73%)方面。各应用间情绪差异显著,其中Claude表现出最高负面情绪(47.7%),同时拥有强烈热情的用户群体,表明存在统计上显著的极化现象。尽管各应用评论数量不等,这些发现依然稳健。作为探索性观察,部分DeepSeek评论提出了与其中国来源相关的地缘政治和数据隐私担忧,而提出的信任摩擦评分(Trust Friction Score)将特定应用的信任和可用性障碍总结为可解释的维度。该研究为消费级生成式AI应用中的用户信任、可用性和采纳障碍提供了经过验证且可操作的证据。

英文摘要

Generative AI (GenAI) applications have achieved rapid consumer adoption, yet little large-scale research examines user-perceived quality, trust, and adoption barriers. We present one of the first cross-application analyses of app store reviews for six major GenAI applications (ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity), comprising 17,012 English-language reviews from Google Play and the Apple App Store. We combine BERTopic topic modeling with RoBERTa sentiment classification and evaluate cross-application differences using chi-square, Kruskal-Wallis, and multinomial logistic regression with Bonferroni correction. Both components are validated against human coding using a stratified sample of 300 reviews. Results show that negative sentiment concentrates in advertising (91%), authentication (89%), server reliability (83%), and subscription pricing (73%). Sentiment differs significantly across applications, with Claude exhibiting the highest negative sentiment (47.7%) alongside a strongly enthusiastic user base, indicating statistically significant polarization. These findings are robust despite unequal review counts across applications. As exploratory observations, a subset of DeepSeek reviews raised geopolitical and data privacy concerns related to its Chinese origin, while a proposed Trust Friction Score summarizes application-specific trust and usability barriers into interpretable dimensions. The study provides validated and actionable evidence on user trust, usability, and adoption barriers in consumer generative AI applications.

CommentsSubmitted to Array (Elsevier); currently under peer review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑