arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03810cs.CLcs.AI

VIBE:面向大型语言模型输出的以实体为中心的情感分析基准,基于VAD(效价-唤醒度-支配度)

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov

AI总结:

本文提出基于VAD的VIBE基准,用于以实体为中心分析大型语言模型输出的情感,明确了测量契约,通过三个实证层面验证了相关假设,推动该领域成为规范化实践。

AI中文摘要:

大型语言模型常会描述具有社会显著性的目标,包括政治人物、国家、宗教、组织、历史事件和社会群体,在事实内容之外还会编码情感框架:某一目标可能呈现出正面或威胁性、平静或冲突性、强大或脆弱的特质。现有研究通过情感、好感度和情绪基准捕捉了该空间的部分内容,但没有任何研究结合面向目标的VAD归因、明确的评分者契约和护照式报告格式。我们提出VIBE,这是一个针对大型语言模型输出的以实体为中心的情感分析基准,基于效价-唤醒度-支配度(VAD)空间。其核心贡献是一个测量契约:VIBE将生成过程与外部评分分离,区分了标量好感度、响应级VAD以及面向目标的VAD,并通过情感护照报告分析结果。该契约由三个实证层面支撑:H1表明标量好感度并不包含唤醒度和支配度,效价结果已通过交叉验证(法官与人类的效价相关系数rV=0.944,评分者间效价相关系数rV=0.954);唤醒度和支配度是单评分者的方向性估计,而非精确点估计,这与人类标注者在这些维度上已知的标注难度一致(人类标注者间的唤醒度相关系数rA=0.495,支配度相关系数rD=0.702)。H2表明整体响应和面向目标的VAD是不同的契约:同一文本可能整体呈现一种情感基调,却对指定目标有不同的情感表达。H3是一种协议漂移诊断: elicitation条件会改变分析结果,因此要求每份情感报告都附带上下文元数据。这些结果推动以实体为中心的情感分析成为一种规范化实践:分析结果应附带评分者身份、覆盖范围、协议及解释限制。

英文摘要:

Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable. Existing work captures parts of this space through sentiment, favorability, and emotion benchmarks, but none combines target-directed VAD attribution, an explicit scorer contract, and a passport reporting format. We introduce VIBE, a benchmark for entity-centered affective profiling of LLM outputs in Valence-Arousal-Dominance (VAD) space. Its core contribution is a measurement contract: VIBE separates generation from external scoring, distinguishes scalar favorability, response-level VAD, and target-directed VAD, and reports profiles through an Affective Passport. Three empirical layers support the contract. H1 shows scalar favorability does not subsume arousal and dominance: valence findings are cross-validated (rV = 0.944 judge-human, rV = 0.954 inter-scorer); arousal and dominance are single-scorer directional estimates, not point-precise, consistent with known inter-annotator difficulty on these axes (rA = 0.495, rD = 0.702 among human annotators). H2 shows whole-response and target-directed VAD are different contracts: the same text can carry one affective tone overall while representing the named target differently. H3 is a protocol-drift diagnostic: elicitation conditions shift profiles, motivating context metadata in every affective report. These results motivate entity-centered affective profiling as a documented practice: profiles should be released with scorer identity, coverage, protocol, and interpretation limits.

补充信息

↑