AI 中文总结
该研究推出ShiJianBench离线框架,用多智能体投资者模拟器评估对话式投资顾问,发现LLM顾问在个性化内容与投资者轨迹上表现更优,凸显需开展轨迹感知评估。
AI 中文摘要
对话式投资顾问不仅影响用户的认知,还会在市场环境演变时影响用户后续的决策方式。现有评估主要衡量响应质量或观测结果,难以审计从顾问语言到投资者行为的长周期路径。我们推出ShiJianBench,这是一个在固定历史市场反馈下通过匹配的投资者轨迹评估对话式投资顾问的离线框架。其核心是多智能体投资者模拟器,具备明确的演化状态变量、动机驱动的 deliberation( deliberation 译为“思考过程”)、长期记忆以及基于对话的更新机制。该模拟器基于7199名真实用户的总体行为模式进行校准,且在严格合规门槛下,通过独立的投资者侧、服务侧和内容侧指标评估顾问策略。对2021至2026年中国基金市场轨迹的实验,识别出一组稳定领先的LLM顾问,它们结合了显著更强的个性化内容与具竞争力的投资者侧轨迹结果。这些结果揭示了生成高质量响应与实施有效长周期干预之间的系统性差异,为对话式顾问的轨迹感知评估提供了动力。
英文摘要
Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving the long-horizon pathway from advisor language to investor behavior difficult to audit. We introduce ShiJianBench, an offline framework for evaluating conversational investment advisors through matched investor trajectories under fixed historical market feedback. At its core is a multi-agent investor simulator with explicit evolving state variables, motive-driven deliberation, long-term memory, and dialogue-grounded updates. The simulator is calibrated against aggregate behavioral patterns from 7,199 real users, and advisor policies are evaluated using separate investor-side, service-side, and content-side metrics under a hard compliance gate. Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes. These results reveal a systematic distinction between producing a high-quality response and delivering an effective long-horizon intervention, motivating trajectory-aware evaluation of conversational advisors.