arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于高效令牌和语义保留观点摘要的大语言模型

Large Language Models for Token-Efficient and Semantic-Preserving Opinion Summarization

Fabrizio Marozzo, Stefano Iannicelli

arXiv 2607.10825首次发表:更新:

发表机构

University of Calabria(卡拉布里亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何在大语言模型进行观点摘要时保留语义并减少令牌使用,结合多维分类与分层抽样策略选择观点子集,定制提示生成摘要,实验证明该方法能降本增效,优于传统和标准大语言模型摘要基线。

AI 中文摘要

带有观点的文本,如产品评论、酒店反馈和社交帖子,包含有关用户体验、偏好和关注点的丰富信号。然而,此类语料库的规模、冗余性和不平衡性使得有效分析观点具有挑战性,尤其是当目标是生成忠实于所表达观点多样性的摘要时。本文提出了一个框架,在基于大语言模型的观点摘要中保留语义,同时尽量减少令牌使用。我们将多维分类(如情感、主题)与一系列分层抽样策略相结合,在提示大语言模型之前选择紧凑而有代表性的观点子集。定制的提示随后生成平衡的摘要,突出观点中表达的显著方面(如产品/酒店的优点和缺点)。在亚马逊产品评论、猫途鹰酒店评论和X/推特帖子上的实验表明,我们的方法显著降低了令牌使用量和计算成本,同时在内容覆盖、平衡和语义保留方面始终优于传统的基于人工智能和标准大语言模型的摘要基线。

英文摘要

Opinionated text - spanning product reviews, hotel feedback, and social posts - captures rich signals about user experiences, preferences, and concerns. However, the scale, redundancy, and imbalance of such corpora make it challenging to analyze opinions effectively, particularly when the goal is to generate summaries that remain faithful to the diversity of viewpoints expressed. This paper presents a framework that preserves semantics in LLM-based opinion summarization while minimizing token usage. We combine multidimensional classification (e.g., sentiment, topics) with a family of stratified sampling strategies to select compact yet representative subsets of opinions before prompting the LLM. Tailored prompts then produce balanced summaries that surface the salient aspects expressed in the opinions (e.g., strengths and weaknesses of products/hotels). Experiments on Amazon product reviews, Tripadvisor hotel reviews, and X/Twitter posts demonstrate that our method significantly reduces token usage and computational cost while consistently outperforming traditional AI-based and standard LLM summarization baselines in terms of content coverage, balance, and semantic preservation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑