新兴零售投资组合管理应用:带有自然语言目标的个性化、税务感知强化学习
An Emerging Retail Portfolio Management Application: Personalized, Tax-Aware Reinforcement Learning with Natural Language Goals
浏览论文内容
中文总结 AI 辅助
该研究针对零售投资者缺乏个性化税务感知投资组合管理的问题,开发了一款集成FastAPI后端等组件的应用,采用三阶段强化学习系统生成推荐,经集成测试后提供预部署验证。
中文摘要 AI 辅助
零售投资者无法获得机构客户理所当然享有的个性化、税务感知投资组合管理服务——现有机器人顾问采用静态、基于规则的配置,而机构级系统要求账户最低金额和个人投资者无法获取的技术栈。我们推出了一款完全构建、经过集成测试的应用,填补了这一缺口:FastAPI后端和网络仪表板,允许用户用自然语言描述投资目标(例如“我希望稳步增长,但需要在下个月卖出部分股票作为首付”),将该目标路由至六项投资指令之一,并通过一个三阶段强化学习系统生成与券商集成的实时投资组合推荐——该系统包含自监督跨资产编码器、带有学习意图路由器的混合专家(Mixture-of-Experts, MoE)配置策略,以及轻量级LoRA适配器,可基于个人披露的券商行为个性化推荐,无需重新训练共享模型。该系统功能完整,已针对实时券商API(Alpaca,模拟交易模式)进行端到端集成测试,包括多用户身份验证、信任优先的先预览后应用确认流程、每日电子邮件摘要和可审计的操作完整性链,但尚未向真实终端用户开放;我们如实报告这是一款具有完整部署路径的新兴预部署应用,同时提供14天滚动窗口回测(包含自举置信区间)作为预部署验证,而非生产性能。我们还报告了若干实用工程经验——静默失效的集成路径、挂起的第三方API调用,以及端到端经验验证的价值,而非依赖检查点元数据——我们认为这些经验可推广至其他基于外部实时数据源构建的应用强化学习系统。
英文摘要
Retail investors lack access to the kind of personalized, tax-aware portfolio management that institutional clients take for granted -- existing robo-advisors use static, rule-based allocation, and institutional-grade systems require account minimums and technology stacks unavailable to individual investors. We present a fully built, integration-tested application that closes this gap: a FastAPI backend and web dashboard that let a user describe an investment goal in plain language (e.g. "I want steady growth but need to sell some shares next month for a down payment"), routes that goal to one of six investment mandates, and produces a live, broker-integrated portfolio recommendation from athree-phase reinforcement learning system -- a self-supervised cross-asset encoder, a Mixture-of-Experts (MoE) allocation policy with a learned intent router, and a lightweight LoRA adapter that personalizes recommendations from an individual's revealed brokerage behavior without retraining the shared model. The system is functionally complete and integration-tested end-to-end against a live brokerage API (Alpaca, paper-trading mode), including multi-user authentication, a trust first preview-before-apply confirmation flow, daily email digests, and an auditable action-integrity chain, but has not yet been opened to real end-users; we report this honestly as an emerging, pre-deployment application with a concrete path to full deployment, alongside 14-day walk-forward backtests (bootstrapped confidence intervals included) as preliminary, pre-deployment validation rather than production performance. We also report several practical engineering lessons -- silently-inactive integration paths, hanging third-party API calls, and the value of end-to-end empirical verification over trusting checkpoint metadata -- that we believe generalize to other applied RL systems built on external, live data sources.