arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20707cs.IR

面向人类购物行为的忠实模拟

Towards Faithful Simulation of Human Shopping Behavior

Jiakai Tang, Yan Mi, Jing Yu, Yang Zhang, See-Kiong Ng, Qi Cao, Fei Sun, Xu Chen, Wen Chen, Jian Wu, Han Zhu, Bo Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有用户购物行为模拟器的记忆与优化挑战,提出分层记忆的RecVerse智能体,结合轨迹级RL优化,发布USB数据集,在模拟性能上显著优于现有基线。

中文摘要 AI 辅助

模拟真实用户购物行为是电商场景下离线评估与强化学习的基础。尽管近期基于LLM(大语言模型)和VLM(多模态大语言模型)的模拟器已取得令人鼓舞的进展,但重现真实浏览会话仍面临两大挑战:(i)记忆挑战:一次购物会话跨越数十个页面,而现有智能体要么丢弃长程观测历史,丢失不断演化的用户状态,要么简单拼接这些历史,导致上下文窗口过载甚至降低模拟质量;(ii)优化挑战:当前用户模拟器通常通过模仿学习或步级奖励监督匹配每条记录的动作,生成的会话常呈现不现实模式,如过度探索或过度被动,而逐步监督既无法检测也无法纠正这些问题。为应对上述挑战,本文提出RecVerse——一种基于GUI(图形用户界面)的模拟智能体,通过截图感知页面并生成忠实的多轮轨迹。针对记忆挑战,RecVerse采用受认知启发的分层记忆:用于短期聚焦的工作记忆、用于会话内轨迹的情景记忆、用于高层意图的偏好记忆,记忆更新被视为动作,使智能体自适应学习何时记忆、记忆什么。针对优化挑战,RecVerse采用轨迹级RL(强化学习)目标优化,对整个会话评分,使宏观动作类型分布与微观购物意图均与真实用户对齐。本文还发布了USB(用户模拟基准)——一个用于多轮用户模拟的交互式电商GUI轨迹数据集。实验表明,RecVerse在行为保真度和意图一致性方面均显著优于现有基线。

英文摘要

Simulating realistic user shopping behavior underpins offline evaluation and reinforcement learning in e-commerce scenarios. While recent LLM- and VLM-based simulators have made encouraging progress, reproducing a real browsing session remains difficult for two reasons. (i) Memory Challenge: a shopping session spans dozens of pages, yet existing agents either discard long-range observation histories, losing the evolving user state, or naively concatenate them, overwhelming the context window and even degrading simulation quality. (ii) Optimization Challenge: current user simulators are typically supervised to match each logged action via imitation or step-level rewards; the resulting sessions often display unrealistic patterns, such as over-exploration or excessive passivity, which per-step supervision can neither detect nor correct. To address the above challenges, we present RecVerse, a GUI-grounded simulation agent that perceives pages through screenshots and produces faithful multi-turn trajectories. For the memory challenge, RecVerse adopts a cognitive-inspired hierarchical memory: Working Memory for short-term focus, Episodic Memory for in-session traces, and Preference Memory for high-level intent, with memory updates treated as actions so that the agent adaptively learns when and what to memorize. For the optimization challenge, RecVerse is optimized with a trajectory-level RL objective that scores entire sessions, aligning both macro-level action-type distributions and micro-level shopping intent with real users. We further release USB (User Simulation Benchmark), an interactive e-commerce GUI trajectory dataset for multi-turn user simulation. Experiments show that RecVerse significantly outperforms existing baselines in both behavioral fidelity and intent consistency.

↑