AI 中文总结
研究针对大语言模型人格分类可解释性问题,提出LEX-EC框架,结合多种诊断与词汇消融,通过对不同文本体裁分析,展示特质关联与词汇证据关系,联合评估多项指标,将词汇方法应用于黑盒人格标签可解释性。
AI 中文摘要
大语言模型能轻松从文本中分配人格标签,但模型可解释性仍是未解决问题。为填补这一空白,我们引入LEX-EC,这是一个可重复使用的黑盒审计框架,结合流行度和一致性诊断与受控词汇消融,以区分边际分布效应与在受限证据下可恢复的特质相关信号。利用该框架,我们展示了不同文本体裁表现出截然不同的特征。自由形式的论文文本包含最广泛但仍微弱的信号;研究生介绍中,外向性关联在屏蔽后减弱;单个脸书状态即使在特质平衡样本中也几乎没有稳定证据。屏蔽主题和人口统计内容削弱了一些关联,而其他关联可从功能词、情感词和认知风格词汇中检测到。语言提示改变了模型的自我解释,但未消除主题内容。LEX-EC联合评估分类流行度、项目级关联、机会校正一致性、词汇限制下的持久性以及模型生成解释中的提示敏感性。在跨数据集、模型和提示中,LEX-EC描述了特质关联如何随可用词汇证据变化,将词汇方法引入人格标签黑盒可解释性的新应用。
英文摘要
Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates classification prevalence, item-level association, chance-corrected agreement, persistence under lexical restriction, and prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling.
CommentsAppendix and link to Code repo provided; this version also contains a refined Discussion section and a small error regarding Table 1 was corrected