arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

引出用于强化学习投资组合优化的ESG偏好

Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization

Giovanni Dispoto, Marcello Restelli, Carmine Ventre

arXiv 2609.02677首次发表:更新:

发表机构

Politecnico di Milano; King’s College London(米兰理工大学; 伦敦国王学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将ESG感知投资组合优化转化为多目标强化学习问题,整合三家ESG机构评级,通过高斯过程偏好引出框架结合LLM角色模拟,发现区域背景影响ESG与收益偏好权重,提供适配现实偏好的算法交易框架。

AI 中文摘要

现代投资组合管理日益需要在传统风险调整后收益与严格的环境、社会和治理(ESG)要求之间取得平衡。当前的强化学习(RL)方法通常针对单一ESG提供商进行优化,忽略了行业内评级方法的显著差异,以及手动对冲突目标进行加权的不直观性。本文通过将ESG感知投资组合优化表述为多目标强化学习(MORL)问题,同时纳入三家不同ESG机构的评级,来解决这些局限性。为弥合高维算法权衡与人类决策之间的差距,我们整合了基于高斯过程的偏好引出框架。该系统使从业者能够通过对候选投资组合的夏普比率和综合ESG评分进行直观的成对比较,推断其潜在效用函数。我们通过使用大语言模型(LLM)角色模拟在不同区域背景下运作的投资组合经理,对我们的框架进行系统评估。使用历史市场数据的实证结果表明,区域背景从根本上改变了推导的偏好权重。例如,基于欧洲的角色倾向于将ESG一致性置于财务收益之上,而基于得克萨斯州的角色则倾向于风险调整后表现。这项工作提供了一个高度适应性的框架,成功地将多目标算法交易与多样化的、现实世界的人类可持续发展偏好相匹配。

英文摘要

Modern portfolio management increasingly demands a balance between traditional risk-adjusted returns and strict Environmental, Social, and Governance (ESG) mandates. Current Reinforcement Learning (RL) approaches typically optimize for a single ESG provider, neglecting the significant divergence in rating methodologies across the industry and the unintuitive nature of manually weighting conflicting objectives. This paper addresses these limitations by formulating ESG-aware portfolio optimization as a Multi-Objective Reinforcement Learning (MORL) problem that simultaneously incorporates ratings from three distinct ESG agencies. To bridge the gap between high-dimensional algorithmic trade-offs and human decision-making, we integrate a Preference Elicitation framework using Gaussian Processes. This system enables practitioners to infer their latent utility functions through intuitive pairwise comparisons of candidate portfolios based on their Sharpe ratios and aggregate ESG scores. We systematically evaluate our framework by employing Large Language Model (LLM) personas to simulate Portfolio Managers operating under varied regional contexts. Empirical results using historical market data reveal that regional backgrounds fundamentally shift the derived preference weights. For instance, European-based personas tend to prioritize ESG alignment over financial returns, while Texas-based personas favor risk-adjusted performance. This work offers a highly adaptable framework that successfully aligns multi-objective algorithmic trading with diverse, real-world human sustainability preferences.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑