自主研究项目管理作为智能体技能:精确谱空间回归的案例研究
Autonomous Research Project Management as an Agent Skill: A Case Study in Exact Spectral Spatial Regression
浏览论文内容
中文总结 AI 辅助
本研究展示了一个智能体技能在消费级硬件上自主执行机器学习研究,通过FFT核岭回归求解器在NOAA数据上验证,实现了闭环科学韧性和最少人类干预,并强调了可检查状态与透明报告的重要性。
中文摘要 AI 辅助
本工作展示了由智能体技能在消费级硬件上自主进行机器学习研究的端到端演示。该演示评估了一种基于FFT的核岭回归(KRR)求解器,用于规则空间网格,使用了2005年月度NOAA Kaplan SST v2异常场,网格大小为$36 \ imes 72$。该实验由DeepSeek V4 Flash自主执行,由我们在DeepSeek Harness(DSH)中的智能体技能套件编排。实验仅在CPU硬件(Apple M2 Pro;求解器时间78.7秒,峰值RSS 1.57 GB)上执行。长时程状态被解耦为基于文件的史诗和问题跟踪底层。在74个子智能体会话中,智能体展示了闭环的科学韧性:将两个失败的假设审查门路由回文献检索,修补了引导索引错误,并且仅通过四次离散的人类引导事件执行。最后,我们反思了自主研究治理,认为科学可信度需要可检查的状态、可证伪的审查门和负面结果的透明报告,并敦促机器学习社区倾向于智能体可访问的结构化格式,而非静态PDF手稿。
英文摘要
This work presents an end-to-end demonstration of autonomous machine learning research conducted by an agent skill on consumer hardware. The demonstration evaluates an FFT-based Kernel Ridge Regression (KRR) solver for regular spatial grids using 2005 monthly NOAA Kaplan SST v2 anomaly fields on a $36 \times 72$ grid. This was autonomously executed by DeepSeek V4 Flash, orchestrated by our agent skill suite within DeepSeek Harness (DSH). Experiments were executed on CPU-only hardware (Apple M2 Pro; 78.7 s solver time, 1.57 GB peak RSS). Long-horizon state was decoupled into a file-based epic- and issue-tracking substrate. Across 74 sub-agent sessions, the agent demonstrated closed-loop scientific resilience: routing two failed hypothesis review gates back to literature retrieval, patching bootstrap indexing bugs, and executing with only four discrete human steering events. Finally, we reflect on autonomous research governance, arguing that scientific credibility requires inspectable state, falsifiable review gates, and transparent reporting of negative results, urging the machine learning community to favour agent-accessible structured formats over static PDF manuscripts.