arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniK:面向数字与物理AI的通用知识感知

UniK: Universal Knowledge Perception for Digital and Physical AI

Nirmit Desai, Kunal Sawarkar, Aditya Mahakali, Dongkon Lee, Kevin Park, Eric Song

arXiv 2609.23971首次发表:更新:

发表机构

AIntropy AI(AIntropy AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UniK提出通用知识感知平台,通过多索引融合检索,无需微调即可在数字与物理AI领域超越大规模专有模型,解决异构知识获取瓶颈。

AI 中文摘要

两类变革性的AI系统正在重塑组织的运作方式:\textit{数字AI},它基于企业知识进行推理,为聊天机器人和智能体工作流提供动力;以及\textit{物理AI},它从视频、游戏玩法和传感器遥测中学习控制机器人和自主系统。两者面临相同的基础性瓶颈:大规模原始知识,跨越异构模态,锁定在现有AI基础设施无法可靠或高效访问的私有语料库中。我们提出\textit{通用知识感知(UniK)}作为这两类系统的通用平台,覆盖从富文本、视频到分子数据和传感器遥测等模态的完整知识生命周期(摄取、丰富、索引、检索和持续评估)。我们展示了UniK,它基于Polymath Retrieval(对自动丰富索引的多索引融合),无需任务特定微调。在五个数字AI领域(医学文献、开放域问答、化学、法律视频诉讼和政府开放数据)中,UniK与一个开源的700亿参数模型相结合,持续匹配或超越规模大数个数量级的专有前沿LLM:政府数据上的RAG准确率为76%,而GPT-5为47%;医学问答上无需微调即达77.9%;在所有开源化学流程中名列前茅。我们表明,相同的基础设施直接解决了物理AI世界模型训练面临的数据整理、索引和检索挑战,其中知识问题更难但结构上相同。

英文摘要

Two transformative classes of AI systems are reshaping how organizations operate: \textit{digital AI}, which reasons over enterprise knowledge to power chatbots and agent workflows; and \textit{physical AI}, which learns to control robots and autonomous systems from video, gameplay, and sensor telemetry. Both face the same foundational bottleneck: raw knowledge at scale, spanning heterogeneous modalities, locked in private corpora that existing AI infrastructure cannot access reliably or efficiently. We propose \textit{Universal Knowledge Perception (UniK)} as a common platform for both classes, covering the full knowledge lifecycle (ingestion, enrichment, indexing, retrieval, and continuous evaluation) across modalities from rich text and video to molecular data and sensor telemetry. We present UniK, built on Polymath Retrieval (multi-index fusion over automatically enriched indices) with no task-specific fine-tuning. Across five digital AI domains (medical literature, open-domain QA, chemistry, legal video proceedings, and government open data) UniK combined with an open-source 70-billion-parameter model consistently matches or outperforms frontier proprietary LLMs that are orders of magnitude larger: 76\% RAG accuracy on government data versus 47\% for GPT-5; 77.9\% on medical QA without fine-tuning; topping all open-source chemistry pipelines. We show that the same infrastructure directly addresses the data curation, indexing, and retrieval challenges facing physical AI world model training, where the knowledge problem is harder but structurally identical.

Comments17 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑