arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27165cs.CLcs.AI

计数证据而非句子:面向长文本价值测量的LLM判断温度化证据融合

Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement

  • HKUST(GZ)(香港科技大学(广州))
  • FDU(复旦大学)
  • DUFE(东北财经大学)
  • UESTC(电子科技大学)
  • UMD(马里兰大学)
  • SEU(东南大学)
  • USTC(中国科学技术大学)
  • Georgia Tech(佐治亚理工学院)
  • CUHK(SZ)(香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

Yuhe Wu, Rui Qian, Guangyu Wang, Yuran Chen, Yuanchao Zhu, Junjie Yang, Zhengheng Li, Jiulin Cai, Tianyi Zhang, Zihan Dong, Jiaxin Liu, Yujie Chen, Guang Zhang

中文总结 AI 辅助

针对长文本价值测量中LLM判断融合的过自信与信息不均问题,提出无需训练的温度化证据融合(TEF),按信息增益加权句子证据,并在新基准MIND上平均提升4.5%准确率与4.6%宏F1。

中文摘要 AI 辅助

大型语言模型(LLM)越来越多地被用于从长篇社交媒体帖子中测量公众价值取向,然而这类帖子往往混杂着背景信息、引述、让步以及少量承载立场的句子。现有方法要么直接让模型预测文档级标签,这可能导致过度自信;要么通过多数投票或软投票聚合句子级预测,将不确定的句子与决定性的句子视为同等信息量。我们将长文本价值测量表述为一个决策融合问题,并提出温度化证据融合(TEF),这是一种无需训练的规则,它根据从广义贝叶斯后验推导出的归一化信息增益对每个句子的对数几率进行加权。这使得不确定句子的融合分数几乎消失,同时保留决定性证据的贝叶斯最优权重。我们进一步引入了多事件洞察网络维度(MIND),这是一个包含8,358条中英文帖子的基准,涵盖五年间的公共事件和六个价值维度。在MIND上,TEF在五个LLM和两种语言中,相比最强的基线(直接预测、多数投票和软投票),平均准确率提升4.5个百分点,宏F1提升4.6个百分点。MIND数据集和代码可在https://github.com/Kzczc/ICASSP2027-TEF获取。

英文摘要

Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be overconfident, or aggregate sentence-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative. We formulate long-text value measurement as a decision-fusion problem and propose Tempered Evidence Fusion (TEF), a training-free rule that weights each sentence's log-odds by its normalized information gain, as derived from a generalized Bayesian posterior. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes-optimal weight of decisive evidence. We further introduce Multi-event Insight Network Dimensions (MIND), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4.5 accuracy points and 4.6 macro-F1 points across five LLMs and two languages. MIND dataset and code are available at https://github.com/Kzczc/ICASSP2027-TEF.

补充信息

↑