从大语言模型激活值中测量文本的概念内容:基于概念向量与线性探测的ESG证据
Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes
浏览论文内容
中文总结 AI 辅助
该研究提出用冻结LLMs的激活值测量文本概念内容,在ESG金融文本实验中,线性探测效果接近微调分类器,优于模型自身答案,且优于RFM概念向量。
中文摘要 AI 辅助
现有测量文本关于某一概念程度的方法仅关注文本表面:词典词频、主题比例、嵌入相似度,它们评分的是文本所用词语,而非读者对文本形成的判断。近期研究表明,大语言模型(LLMs)内部知识与它们在响应中表达的内容存在差距。本文探究通过监测冻结的、开箱即用的LLMs激活值所读取的内部知识,能否在测量概念内容时替代特定任务微调,以及哪种提取方法效果最佳。我们通过递归特征机(RFM)算法和线性探测提取此类测量值,并将其与嵌入基线、表面基线以及同一模型对问题的自身答案进行比较。我们在金融文本(该领域已被广泛研究且拥有成熟标注资源)上展示了该方法,使用人工标注的环境、社会和治理(ESG)数据集。最佳线性探测的准确率与微调后的领域分类器仅相差0.6个百分点,且在12次比较中有11次超过同一模型对问题的自身答案,因此激活值承载了响应未报告的概念内容。简单探测始终优于RFM概念向量,而RFM概念向量提供了仅分类无法做到的内容:一个旨在反映文本中概念存在强度的连续分数,其验证有待分级标签。
英文摘要
Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities. They score the words a text uses, not the judgment a reader forms about it. Recent work has shown that a gap exists in what Large Language Models (LLMs) know internally versus what they express in their response. This paper asks whether that internal knowledge, read by monitoring the activations of frozen, out-of-the-box LLMs, can stand in for task-specific fine-tuning when measuring concept content, and which extraction method reads it best. We extract such measures via the Recursive Feature Machine (RFM) algorithm and via linear probing, and compare these against an embedding baseline, surface baselines, and the same model's own answer to the question. We demonstrate the approach on financial text, a domain studied extensively and served by established annotated resources, using a human-annotated Environmental, Social and Governance (ESG) dataset. The best linear probe comes within 0.6 percentage points of a fine-tuned domain classifier's accuracy without any task-specific fine-tuning, and outscores the same model's own answer to the question in eleven of twelve comparisons, so the activations carry concept content the response does not report. The simple probe consistently beats the RFM concept vectors, which in turn provide what classification alone does not: a continuous score intended to reflect how strongly a concept is present in a text, whose validation awaits graded labels.
发表机构
- Leiden Institute of Advanced Computer Science, Leiden University(莱顿大学高级计算机科学研究所)
机构由 AI 辅助整理,请以论文原文为准。