arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型中风险规避的分布外泛化

Out-of-Distribution Generalization of Risk Aversion in Language Models

Kristina Zhang, Junior Chinomso Okoroafor, Benjamin Maltbie, Andrew Lin, Abhitej Bokka, Elliott Thornley

arXiv 2607.02755首次发表:更新:

AI 中文总结

研究训练人工智能在资源方面规避风险能否泛化至高风险场景。通过引入RiskAverseOOD基准测试,用多种方法使模型在低风险时规避风险,发现其能部分泛化至高风险,不同规模和模型家族也有类似效果,但一致性不足。

AI 中文摘要

训练人工智能在资源方面规避风险,若人工智能出现对齐问题,这可提供一个故障安全机制。但只能在低风险赌博上训练其规避风险,只有当这种风险规避能泛化到极高风险赌博时才安全。为此引入RiskAverseOOD基准测试并给出初步结果,发现低风险时学到的风险规避能部分泛化到极高风险,不同规模和模型家族也有类似情况,但一致性尚不足以作为可靠的故障安全机制。

英文摘要

Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AIs would tend to prefer low-risk, low-reward strategies like cooperation over high-risk, high-reward strategies like rebellion, limiting the downsides of any misalignment. But we can only feasibly train AIs to be risk-averse on low-stakes gambles, and we will only be safe if their risk aversion generalizes to astronomically-high-stakes gambles. Will it? To shed light on this question, we introduce RiskAverseOOD: a benchmark for measuring how well risk aversion generalizes out of distribution. We then offer some initial results. Using a variety of methods to make Qwen3-8B choose risk-aversely when the stakes are low, we find that we can induce substantial risk aversion when the stakes are astronomically high. From a baseline 2% rate of choosing a safe `Cooperate' option, we see rates around 70% (SFT and tie training) and 52% (DPO). Activation steering scores 78% but hurts capabilities and makes the model excessively risk-averse. We observe similar effects at different scales (Qwen3-1.7B and Qwen3-14B) and across model families (Gemma-3-12B-IT and Llama-3.1-8B-Instruct). Overall, we find that risk aversion learned at low stakes can generalize OOD to astronomically high stakes, though not yet consistently enough to serve as a reliable failsafe. Achieving that level of consistency is an open problem.

CommentsICML 2026 Agents in the Wild workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑