AI 中文总结
针对外部 API 无法读取代码的问题,提出 PAU 基准,要求模型通过交互查询理解代码功能;发现前沿模型过度自信,并借鉴 AAC 范式后训练,使 Qwen3-8B 达到 GPT-5-mini 水平。
AI 中文摘要
在语言模型智能体取得进展的推动下,系统在代码生成和理解方面取得了长足进步。然而,这些方法通常依赖于对相关代码的读取访问,而这一假设在处理外部 API 时并不成立。在这项工作中,我们引入了 PAU(Python API 理解)基准,在该基准中,我们为模型提供对代码片段的黑盒、API 级访问。模型必须使用探索性输入查询 API,并从由此产生的输出中获取洞见,其目标是描述代码片段的真实功能。通过将代码片段视为必须仅通过交互来理解的外部工具,PAU 研究了无监督工具理解这一更普遍的问题,特别是针对以 Python 方法实现的工具。尽管编码智能体近期取得了进展,但即使是最前沿的模型也难以在 PAU 上取得高性能,最佳模型(Claude-4-Opus)未能理解 PAU 测试集中超过 45% 的内容。对常见错误模式的研究表明,模型过于自信;它们常常高估当前假设的质量,导致探索不足和过早终止。最后,我们借鉴了机器人学习中常用的非对称 Actor-Critic(AAC)范式,对模型进行交互式代码理解的后训练。使用 AAC 训练的模型对 API 进行了更主动的探索,经过 AAC 调优的 Qwen3-8B 模型达到了与 GPT-5-mini 相当的性能。
英文摘要
Empowered by advances in Language Model agents, systems have made substantial strides in code generation and understanding. However, these approaches often rely on read access to the relevant code, an assumption which does not hold when dealing with external APIs. In this work, we introduce the PAU (Python API Understanding) benchmark, where we provide models with black-box, API-level access to code snippets. Models must query the API with exploratory inputs and draw insights from the resulting outputs, with the goal of describing the snippet's true functionality. By treating the code snippets as external tools that must be understood via interaction alone, PAU studies the more general problem of unsupervised tool understanding, specifically for tools implemented as Python methods. Despite recent progress in coding agents, even frontier models struggle to achieve high performance on PAU, with the best model (Claude-4-Opus) failing to understand over 45% of the PAU test set. An investigation into the common error modes reveals that models are overconfident; they often overrate the quality of their current hypothesis, leading to insufficient exploration and premature termination. Finally, we take inspiration from the Asymmetric Actor Critic (AAC) paradigm, frequently used in robot learning, to post-train models for interactive code understanding. Models trained with AAC conduct more active exploration of the APIs, with an AAC-tuned Qwen3-8B model matching the performance of GPT-5-mini.