AgentBetta:通过选择性扩展与验证性收缩实现AI纳米智能体的验证驱动自适应配置
AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction
浏览论文内容
中文总结 AI 辅助
AgentBetta通过验证驱动诊断、选择性扩展和验证性收缩自适应配置AI纳米智能体,在AB-ConfigBench上实现91.38%验证成功并大幅降低资源暴露,作为配置调节机制而非通用替代方案。
中文摘要 AI 辅助
大型语言模型智能体通常以预定义配置部署,尽管不同任务对所需模型能力、上下文、工具、权限、记忆和计算资源的需求差异显著。本研究开发并评估了AgentBetta,一个自适应AI纳米智能体框架,该框架将这些因素表示为可执行配置,并通过验证驱动的诊断、选择性扩展和基于验证的反事实收缩来更新配置。评估区分了受控机制验证与外部智能体比较。在AB-ConfigBench基准上,AgentBetta实现了91.38%的验证成功,同时与完全配置相比,将中位上下文分配从64,000个上下文字符减少至8,000个,中位工具暴露从五个工具减少至零。配置缺陷诊断在评估维度上实现了0.819的宏F1分数,精确率为1.000,选择性扩展避免了对无关配置维度的不必要更改。成功后的收缩在56.41%的评估单维度收缩探针中保持了验证结果,表明某些成功配置在测试条件下包含可移除的能力。外部评估表明,自适应配置可以改善验证任务完成与能力暴露之间的平衡;然而,结果在不同基准和智能体家族间存在差异。特别是,跨家族复制未重现主骨干准确性排序,且专门系统在某些任务领域仍具优势。这些结果支持将AgentBetta解释为一种配置适应机制,用于调节能力分配和推理支出,而非作为专门智能体架构的通用替代品。
英文摘要
Large language model agents are typically deployed with predefined configurations, although the required model capability, context, tools, permissions, memory, and computational resources can vary substantially across tasks. This study develops and evaluates AgentBetta, an adaptive AI Nano-Agent framework that represents these factors as an executable configuration and updates them through verification-driven diagnosis, selective expansion, and verification-based counterfactual contraction. The evaluation distinguishes controlled mechanism validation from external agent comparisons. On the AB-ConfigBench benchmark, AgentBetta achieved 91.38% verified success while reducing median context allocation from 64,000 to 8,000 context characters and median tool exposure from five tools to zero compared with the fully provisioned configuration. The configuration-deficiency diagnosis achieved a macro-F1 score of 0.819 with precision of 1.000 across the evaluated dimensions, and selective expansion avoided unnecessary changes to unrelated configuration dimensions. Post-success contraction preserved verification outcomes in 56.41% of evaluated one-dimension contraction probes, indicating that some successful configurations contained removable capability under the tested conditions. External evaluations indicate that adaptive configuration can improve the balance between verified task completion and capability exposure; however, the results vary across benchmarks and agent families. In particular, the cross-family replication did not reproduce the primary-backbone accuracy ordering, and specialized systems remained advantageous for certain task domains. These results support interpreting AgentBetta as a configuration-adaptation mechanism that regulates capability allocation and inference expenditure rather than as a universal replacement for specialized agent architectures.
发表机构
- Independent University, Bangladesh(孟加拉国独立大学)
机构由 AI 辅助整理,请以论文原文为准。