arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31847cs.CL

Omni-IO Skills:发挥你的智能体全模态原生能力

Omni-IO Skills: Harnessing Your Agent Omni-Native

Yanlin Li, Mingyang Hao, Shengqiong Wu, Hao Fei, Mong-Li Lee, Wynne Hsu

首次发表
浏览论文内容

中文总结 AI 辅助

提出Omni-IO Skills框架,通过层级技能和声明式执行图统一多模态任务,将GPT-5.6 Sol和Claude Sonnet 5的输入支持率提升至100%,显著提高语义-质量耦合分数。

中文摘要 AI 辅助

通用智能体能够进行长时程的规划、推理和行动,但其生产能力仍分散在文本、图像、音频、视频、文档、3D资产和代码等模态中。将基础模型扩展到更多模态,会使能力增长依赖于昂贵的模型更新,而组装专家模型和工具则留下了程序、依赖关系、中间资产和跨轮次修订如何协调的问题。我们提出Omni-IO Skills,一个即插即用的智能体框架(Agent Harness),通过层级化技能(Skills)、标准化的多模态执行接口、依赖感知的编排以及持久化的资产注册表(Asset Registry),使现有智能体具备全模态原生能力。多资产工作流被表示为声明式执行图(Declare Execution Graphs),该图可并发调度独立操作,并将成功输出注册以供下游和跨轮次复用,且可在可替换的执行后端上运行。其27项技能覆盖38个代表性任务,横跨七种工件模态和四类能力家族:理解、生成、推理和检索。在UniM-90基准上,该框架将GPT-5.6 Sol和Claude Sonnet 5的输入支持率从40.00%和38.89%提升至100%,同时将相对语义-质量耦合分数分别从26.99提高到74.94、从27.82提高到77.78;严格结构分数达到100.00和99.78。这些结果表明,在不改变宿主智能体推理核心的前提下,框架级能力组合是实现广泛、可演进的全模态系统的实用途径。

英文摘要

General-purpose agents can plan, reason, and act over long horizons, yet their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets, and code. Extending a foundation model to additional modalities ties capability growth to costly model updates, while assembling specialist models and tools leaves unresolved how procedures, dependencies, intermediate assets, and cross-turn revisions should be coordinated. We present Omni-IO Skills, a plug-and-play Agent Harness that makes existing agents omni-native through hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry. Multi-asset workflows are represented as Declare Execution Graphs, which schedule independent operations concurrently and register successful outputs for downstream and cross-turn reuse across replaceable execution backends. Its 27 Skills cover 38 representative tasks spanning seven artifact modalities and four capability families: understanding, generation, reasoning, and retrieval. On UniM-90, the harness raises the input-support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% to 100%, while increasing relative Semantic--Quality Coupled Score from 26.99 to 74.94 and from 27.82 to 77.78, respectively; Strict Structure Score reaches 100.00 and 99.78. These results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.

补充信息

↑