arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16436cs.LGcs.AIcs.CL

解读与引导用于社会模拟的LLM智能体

Interpreting and Steering LLM Agents for Social Simulations

  • Data Innovation and AI Lab (DIAL), University of California, Berkeley(加州大学伯克利分校数据创新与人工智能实验室)
  • National Bureau of Economic Research (NBER)(美国国家经济研究局)

机构由 AI 辅助整理,请以论文原文为准。

Jiayue Gaveal Fan, Arul Murugan, Shreyas Krishnan, Abhishek Nagaraj

AI总结:

本研究通过比较提示操控、SAE特征引导和探针方向引导三种方法,解读并引导LLM智能体的偏好与能力,发现SAE和探针方法在引导上更有效,为社会科学模拟提供了可解释且可引导的流程。

AI中文摘要:

基于大型语言模型(LLM)的模拟已被证明是理解人类行为的强大工具,使其成为社会科学工具包中的宝贵补充。然而,LLM本质上是基于深度神经网络的黑箱,这限制了其对社会科学的价值。这是因为缺乏(i)可解释性:即为观察到的行为分配清晰机制的能力;以及缺乏(ii)可引导性:即放大或抑制特定理论上有意义的行动机制以驱动特定模型行为的能力。在此,我们展示了如何打开黑箱以进一步丰富基于LLM的模拟。具体而言,我们比较了三种方法:(1)基于提示的操控,(2)源自SAE的特征引导,以及(3)基于探针的方向引导,并考察了它们在基于LLM的社会科学模拟中的效用。我们通过解读和引导人类行为的两个基础组成部分来实现这一点,即偏好(风险态度、利他主义)和能力(发散性创造力、产品创新),这些通过四个经典的经济和创造性任务以自然语言交互的形式进行操作化。总体而言,我们的结果表明,基于SAE和探针的技术在引导LLM智能体方面通常优于基本的基于提示的方法,尽管这种优势取决于所涉及的特定提示策略。总之,SAE和探针构成了一个有效的流程,供寻求在社会模拟中解读和引导智能体的社会科学家使用:SAE将智能体的内部表征分解为人类可读的特征,之后探针可以可靠地将智能体的行为转向指定方向。我们讨论了这些方法对未来使用LLM智能体进行社会科学模拟工作的意义。

英文摘要:

Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions. Overall, our results show that SAE- and probe-based techniques often outperform basic prompt-based methods for steering LLM agents, although this advantage depends on the specific prompting strategy involved. Together, SAEs and probes constitute an effective pipeline for social scientists seeking to interpret and steer agents in social simulations: SAEs decompose agents' internal representations into human-readable features, after which probes can reliably shift agents' behaviors in specified directions. We discuss implications of these methods for future work using LLM agents for social scientific simulations.

补充信息

↑