AI 中文总结
本文针对主动式AI助手个性化交互时机难确定的问题,提出EOPA方法,在ProPerSim基准上较最强基线提升F1分数19.80个百分点,降低推理延迟与平均每日适应时间。
AI 中文摘要
AI助手通常是反应式的,依赖用户发起交互;主动式助手则突破这一范式,可基于用户的活动情境自主发起交互。然而,合适的交互时机具有用户特异性,难以提前确定,而在线反馈为个性化提供了有价值的信号。直接基于反馈的适应颇具吸引力,但由于值得交互的时刻分散在细粒度用户状态中,这一过程仍具挑战性。为解决上述问题,本文提出证据驱动的在线偏好适应(Evidence-driven Online Preference Adaptation, EOPA),该方法通过两种证据载体将用户的交互时机偏好建立在可测量的情境证据上:时间偏好锚点和含证据的活动原型。在每个轮询步骤中,EOPA通过用户先验平滑的证据估计和不确定性引导的证据缩放,从载体中推导时间和活动证据,并自适应融合这些证据以做出交互或弃权(不执行)的决策。当选择交互时,大型语言模型(LLM)会使用高质量历史响应作为演示,生成更贴合用户偏好的情境感知响应。EOPA会基于接收到的在线反馈更新其证据载体和决策参数,无需基于LLM的推理或重新训练。在基于ProPerSim的基准测试上进行的大量实验表明,EOPA将交互时机F1分数较实验中最强基线提升了19.80个百分点,大幅降低了沉默和交互步骤的推理延迟,并将平均每日适应时间从11.41秒降至0.39秒。
英文摘要
AI assistants are typically reactive, relying on users to initiate interactions. Proactive assistants go beyond this paradigm by autonomously initiating interactions based on users' activity contexts. However, appropriate interaction timing is user-specific and difficult to determine in advance, while online feedback offers valuable signals for personalization. Direct feedback-driven adaptation is therefore appealing, but remains challenging due to sparse interaction-worthy moments scattered across fine-grained user states. To address the issues, we propose Evidence-driven Online Preference Adaptation (EOPA), which grounds a user's interaction-timing preferences in measurable contextual evidence through two evidence carriers: temporal preference anchors and evidence-bearing activity prototypes. At each polling step, EOPA derives temporal and activity evidence from the carriers through user-prior-smoothed evidence estimation and uncertainty-guided evidence scaling, and adaptively fuses the evidence for interaction-or-silence decisions. When interaction is selected, an LLM uses high-quality historical responses as demonstrations to generate a context-aware response that better reflects user preferences. EOPA updates its evidence carriers and decision parameters from received online feedback without LLM-based reasoning or retraining. Extensive experiments on a ProPerSim-based benchmark show that EOPA improves the interaction-timing F1 score by 19.80 points over the strongest baseline in our experiments, substantially reduces inference latency for both silence and interaction steps, and lowers the average daily adaptation time from 11.41 to 0.39 seconds.