arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.00362cs.AIcs.LG

与gpt-oss和谐共处

In harmony with gpt-oss

Borislav Mavrin

更新

AI总结:

研究通过逆向工程gpt-oss-20b的工具调用机制,构建了原生和谐代理框架,实现了对OpenAI发布分数的独立复现,提升了工具调用精度和评分表现。

AI中文摘要:

没有人曾独立复现OpenAI发布的gpt-oss-20b分数,因为原始论文未披露工具和代理框架。我们逆向工程了模型的分布内工具:在无工具定义提示下,gpt-oss仍能以高统计置信度调用训练分布中的工具——这是强先验而非幻觉。我们随后构建了原生和谐代理框架(https://github.com/borislavmavrin/harmonyagent.git),该框架将信息编码为模型原生格式,绕过了损失性的Chat Completions转换。这些成果共同实现了对OpenAI发布分数的首次独立复现:在SWE Verified HIGH上达到60.4%(原发布60.7%),在MEDIUM上达到53.3%(原53.2%),在AIME25带工具的情况下达到91.7%(原90.4%)。

英文摘要:

No one has independently reproduced OpenAI's published scores for gpt-oss-20b with tools, because the original paper discloses neither the tools nor the agent harness. We reverse-engineered the model's in-distribution tools: when prompted without tool definitions, gpt-oss still calls tools from its training distribution with high statistical confidence -- a strong prior, not a hallucination. We then built a native harmony agent harness (https://github.com/borislavmavrin/harmonyagent.git) that encodes messages in the model's native format, bypassing the lossy Chat Completions conversion. Together, these yield the first independent reproduction of OpenAI's published scores: 60.4% on SWE Verified HIGH (published 60.7%), 53.3% MEDIUM (53.2%), and 91.7% on AIME25 with tools (90.4%).

↑