arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09684cs.CLcs.AI

从帕累托到偏好:通过摊销智能体策略发现实现个性化测试时扩展

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

Xinglin Wang, Zishen Liu, Tong Zheng, Shaoxiong Feng, Peiwen Yuan, Yiwei Li, Jiayi Shi, Yueqi Zhang, Chuyi Tan, Ji Zhang, Boyuan Pan, Kan Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出PersonTTS框架,通过摊销智能体策略发现重用搜索经验,实现多维度用户需求下的个性化测试时扩展,在AIME和HMMT上显著提升联合需求满足率。

中文摘要 AI 辅助

测试时扩展(TTS)通过分配额外的推理计算来提升大型语言模型的推理能力。现有改进TTS效率的方法主要针对单一资源维度优化准确性,分别推进准确性-成本或准确性-延迟的帕累托前沿。然而,用户需求是多维的:用户可能同时指定准确性、延迟和推理成本需求,且不同需求可能偏好不同的控制器。我们将个性化测试时扩展定义为发现可执行控制器,以最大化用户特定需求的联合满足率。为降低针对新用户画像重复进行策略发现的开销,我们提出PersonTTS,一种摊销智能体策略发现框架,通过需求匹配的控制器初始化和源蒸馏的程序性指导重用先前搜索经验,同时对每个候选保留目标画像评估。在AIME和HMMT上的实验表明,PersonTTS在未见用户画像和保留问题上,在联合需求满足率方面显著优于强TTS基线。在相同候选评估预算下,跨用户经验重用进一步提升了策略质量,同时大幅降低了发现智能体的时间和成本。

英文摘要

Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.

发表机构

  • Beijing Institute of Technology(北京理工大学)
  • Xiaohongshu Inc(小红书公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑