arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

众多观点,众多大语言模型:将大语言模型与传统机器学习用于开放式调查分析的比较

So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis

Abdullah Akinde, Mariam Akinde, Rasheedat Emiola, Ahmed Akinsola

arXiv 2607.11890首次发表:更新:

AI 中文总结

研究对比多种大语言模型与传统机器学习用于开放式调查分析,在情感分析和主题分类等任务中评估性能,发现LLMs分类准确性更高,但在预测理由及应用类别边界上差异大,凸显了使用LLMs进行定性分析时的权衡并给出实用建议。

AI 中文摘要

开放式调查能提供宝贵见解,但大规模分析极具难度。本研究基于先前用传统机器学习对文本分类的工作,探究不同大语言模型(LLMs)如何理解和分析NSSE开放式调查回复。聚焦于OpenAI的GPT系列、Twitter-roBERTa-base模型和Meta的LLaMA等前沿LLMs,并将其与先前机器学习模型在情感分析和主题分类等任务中的性能作比较。评估模型一致性、分类准确性和推理可解释性。结果显示,当前LLMs在分类准确性上常胜过经典机器学习模型,尤其在理解学生回复中的复杂情绪和主题模式方面。然而,LLMs在预测理由的明确性和一致性以及应用类别边界方面差异很大。这些差异凸显了使用LLMs进行定性分析时的关键权衡:预测能力增强伴随着一致性和可解释性问题。研究结果阐明了在大规模定性研究中使用各种LLMs的利弊,并为寻求平衡自动化和解释严谨性的研究人员提供了实用建议。

英文摘要

Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditional machine learning to classify text ("So Many Responses, So Little Time: A Machine-Learning Approach to Analyzing Open-Ended Survey Data") [1], this study investigates how different large language models (LLMs) understand and analyze NSSE open-ended survey responses. We focus on several cutting-edge LLMSs-OpenAI's GPT series, Twitter-roBERTa-base model, and Meta's LLaMA-and compare their performance to the previous machine learning models in tasks like sentiment analysis and thematic classification. Our research analysis assesses model agreement, classification accuracy, and interpretability of reasoning. The findings reveal that current LLMs routinely beat classic machine learning models in classification accuracy, particularly in understanding complex mood and theme patterns in student replies. While LLMs have superior accuracy, they differ greatly in how explicitly and consistently they justify their predictions and apply category boundaries. These distinctions highlight crucial trade-offs when using LLMs for qualitative analysis: increased predictive strength comes with issues in consistency and explainability. Our findings illustrate the benefits and drawbacks of utilizing various LLMs for large-scale qualitative research, and we provide practical advice for researchers looking to balance automation and interpretive rigor.

DOI:10.54536/ari.v4i1.6586

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑