arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

mimeo:将公开专家语料库编译为智能体技能并测试迁移效果

mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers

Timothy Kassis

arXiv 2609.00453首次发表:更新:

发表机构

K-Dense, Inc.(K-Dense公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

mimeo 是一款开源工具,可编译公开专家语料库生成智能体可加载文件,经测试其在知识获取、角色定位上表现优于普通智能体,且无法证明专家判断的迁移效果。

AI 中文摘要

为智能体提供关于某位知名专家的文件,可提供难以获取的资料、塑造可识别的角色,或改变智能体的决策,这些是不同的主张,本文对每一项都进行了测试。mimeo 是一款开源工具,它能查找某人的公开作品,将提取的每一段引述与缓存的源文本进行比对,并生成一个智能体可加载的文件。八次记录的构建过程平均需要 38 次模型调用;比对过程拒绝了 13.2% 的提取引述。我们使用同一个编码智能体测试框架对四份专家文件进行了测试。知识获取方面的效果最为明显:mimeo 回答了全部 20 个晦涩难懂、大量引用的问题;而无背景知识的条件下回答的问题不超过 10 个。对相同页面进行关键词搜索(BM25)可回答 15-17 个问题,该样本无法明确此差距的来源。角色定位方面显示出一项明确的优势:基于模型记忆生成的角色在 20 个答案中,每个评分者都发现有 1-4 个答案错误陈述了专家已记录的立场;而普通智能体和 mimeo 从未出现此类错误。所有角色在简短的开放式提示下都极易被识别,添加任务材料后,识别率降低了 18-23 个百分点。mimeo 的可识别性并不高于基于记忆生成的角色。判断迁移效果仍未明确,因为所有测试都达到了上限:所有条件下,智能体在工程任务中发现了 94-97% 的植入问题,在 16 个新应用场景中得分在 94-100% 之间。由 AI 评判的“听起来像专家”的评分因评判者而异:四名评判者中有两名更喜欢基于模型刻板印象的答案,而另外两名评判者在相同文本上未发现差异。这警示人们不能依赖单一的 AI 评判者。证据表明,mimeo 可作为关于某人的紧凑、可检查的参考资料,而非其判断已成功迁移的证明。工具包和专家资料可通过此 https URL 获取。

英文摘要

Giving an agent a file about a named expert can supply hard-to-find material, produce a recognizable persona, or change what the agent decides. These are different claims. We test each one. mimeo is an open-source tool that finds a person's public work, checks each extracted quotation against the cached source text, and writes a file an agent can load. Eight logged builds averaged 38 model calls; the check rejects 13.2% of extracted quotations. We tested four expert files with one coding-agent harness. Knowledge access was clearest: mimeo answered all 20 obscure, quotation-heavy questions; no closed-book condition answered more than 10. Keyword search (BM25) over the same pages answered 15-17, a gap this sample cannot resolve. Grounding showed one clear benefit: personas written from model memory misstated a documented position on 1-4 of 20 answers under every grader; the plain agent and mimeo never did. Every persona was easy to spot on short open prompts, and adding task material lowered identification by 18-23 points. mimeo was no more identifiable than a from-memory profile. Judgment transfer remained unresolved because both tests hit their ceiling: every condition found 94-97% of the problems planted in engineering tasks and scored 94-100% on 16 new application scenarios. An AI-judged "sounds like the expert" score changed with the judge: two of four preferred answers based on a model's stereotype, while two found no difference on the same text. That is a caution against relying on a single AI judge. The evidence supports mimeo as a compact, inspectable reference on a person, not as a demonstrated transfer of their judgment. Toolkit and expert profiles: https://github.com/K-Dense-AI/mimeo

Comments42 pages, 6 figures. Toolkit and expert profiles: https://github.com/K-Dense-AI/mimeo

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑