发表机构
McGill University; University of Toronto; University of California, Los Angeles; The Chinese University of Hong Kong; University of Manitoba; Université de Montréal; Mila – Quebec AI Institute; McMaster University; Harvard University; Monash University; The University of Hong Kong; Stanford University; Queen Mary University of London; UE Capital Ltd.; CG Matrix Technology Ltd.(麦吉尔大学; 多伦多大学; 加州大学洛杉矶分校; 香港中文大学; 曼尼托巴大学; 蒙特利尔大学; 米拉-魁北克人工智能研究所; 麦克马斯特大学; 哈佛大学; 莫纳什大学; 香港大学; 斯坦福大学; 伦敦玛丽女王大学; UE资本有限公司; CG矩阵科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FINSKILLOPS提出一种自我进化的多智能体系统,通过有范围的技能补丁和受控生命周期管理,在SEC文件问答中实现部署后改进,提升正确性并减少回归。
AI 中文摘要
金融问答系统通常在部署前通过改进检索、提示或智能体协调来提升性能,此后其可靠性行为便固定不变。在实践中,新的SEC文件问题会反复暴露出在时期、实体、证据使用和计算方面的异构错误。现有的自我改进方法可以将失败转化为新的行为,但对修正应应用于何处或可能破坏哪些先前正确的答案提供的控制有限。因此,我们将部署后改进视为受控的行为维护:反复出现的失败应成为有范围的技能补丁,每个补丁应在不引入回归的情况下赢得部署。我们在FINSKILLOPS中实现了这一观点,这是一个用于SEC文件问答的多智能体系统。FINSKILLOPS从基于证据的、类型化的失败诊断中派生可复用技能,并通过定向验证、受保护案例回归检查、阴性对照以及版本化替换或退役来管理这些技能。在六个金融问答基准上,一个单一的冻结技能注册表在评估系统中实现了最高的判决加权正确性和参考一致性。进化后的技能在我们增强的基准上将正确性从3.70提升到4.55。在另一项为期12轮的操作研究中,33个提议的技能中只有6个被提升,而监控非正确率从20.0%下降到12.5%。这些结果确立了受控的技能范围、准入和生命周期管理作为可靠自我改进的基础。
英文摘要
Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation. Existing self-improvement methods can turn failures into new behaviors, but offer limited control over where a correction should apply or which previously correct answers it may break. We therefore frame post-deployment improvement as controlled behavioral maintenance: recurring failures should become scoped skill patches, and each patch should earn deployment with- out introducing regressions. We instantiate this view in FINSKILLOPS, a multi-agent system for SEC filing QA. FINSKILLOPS derives reusable skills from evidence-grounded, typed failure diagnoses and governs them through targeted validation, protected-case regression checks, negative controls, and versioned replacement or retirement. Across six financial QA benchmarks, a single frozen skill registry achieves the highest verdict-weighted correctness and reference consistency among the evaluated systems. Evolved skills raise correctness from 3.70 to 4.55 on our enhanced benchmark. In a separate 12-round operational study, only six of 33 proposed skills are promoted, while the monitoring non-correct rate falls from 20.0% to 12.5%. These results establish controlled skill scope, admission, and lifecycle management as the foundation for reliable self-improvement.