arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI模型的语义对齐:概念坍缩、检查点动态与跨语言迁移

Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer

Tyler Ashoff, Jordan Rodu

arXiv 2608.01585首次发表:更新:

AI 中文总结

该研究针对语言模型基准测试的难点,提出用拓扑方法将模型高维嵌入空间与可解释基线严格比较,以实现多模态对齐,追踪模型适应并测试跨语言短语理解。

AI 中文摘要

语言模型基准测试是一项困难的任务,仅通过结果推理无法测试模型对语言的概念化能力,且流行的开源基准会迅速被饱和或作为训练数据纳入。测试模型的输出很重要,但通过表征语义结构来补充这些测试,能更深入地了解模型与抽象概念的关联方式。不过高维嵌入空间难以解释,本研究展示了如何使用拓扑方法将这些空间与本体、精心构建的知识图谱等低维可解释基线进行严格比较,这些多模态对齐测试可用于追踪模型适应情况,并测试跨多种语言的短语理解能力。

英文摘要

Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important to test the model's output, but augmenting these tests by characterizing semantic structure gives more insight to how models relate abstract concepts. However, the high dimensional embedding spaces are not easy to interpret. This work demonstrates how topological methods can be used to rigorously compare these spaces to low dimensional and interpretable baselines like ontologies and curated knowledge graphs. These multi-modal alignment tests make it possible to track model adaptations and test phrase understanding across multiple languages.

CommentsCode available at github.com/tylerashoff/persiscope (PyPI: persiscope)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑