AI 中文总结
研究针对用户指令不明确时网络代理需从浏览历史推断上下文的问题,引入PersonaTrail基准及偏好感知上下文记忆框架PACMem,利用浏览轨迹评估代理,实验证明PACMem优于现有基于记忆的基线。
AI 中文摘要
大语言模型的进展使网络代理能自主执行复杂任务。实际中用户常提供不明确指令,需代理从原始浏览历史推断缺失上下文。现有基准无法捕捉这种个性化形式。为此引入PersonaTrail,在受管开放网络环境中对个性化网络代理进行基准测试。利用现实浏览轨迹评估代理推断用户偏好和从过往会话召回信息的能力。还提出偏好感知上下文记忆(PACMem)框架,将原始浏览历史分解为事实记忆和偏好记忆。大量实验表明PACMem在两项任务上均优于现有基于记忆的基线。
英文摘要
Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histories. Existing benchmarks fail to capture this form of personalization, as they either restrict tasks to fully explicit prompts or abstract web interaction history into simplified forms. To bridge this gap, we introduce PersonaTrail, a benchmark for personalized web agents operating in a managed open web environment. By leveraging realistic browsing trajectories as user history, PersonaTrail evaluates an agent's ability to infer user preferences and recall information from past browsing sessions. We further propose Preference-Aware Contextual Memory (PACMem), a framework that decomposes raw browsing histories into two types of structured memory: factual memories that summarize individual sessions and preference memories that distill recurring behavioral patterns. At inference time, the agent retrieves the most relevant entries from these memories to guide personalized navigation. Extensive experiments show that PACMem consistently outperforms existing memory-based baselines on both tasks.