arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向以单元为基元的机器学习:从单元关联事件中学习

Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events

Heyang Gong

arXiv 2608.25118首次发表:更新:

AI 中文总结

该研究提出以单元为机器学习的显式语义基元,定义单元关联事件的学习框架,给出监督学习特例的形式化结果,区分单元敏感与不敏感学习,解决单元身份未明确时的学习问题。

AI 中文摘要

机器学习通常以样本为形式化基础,而多个已观测或可能事件所指向的持久个体往往隐含未明。我们提出将「单元(unit)」作为任务语义层面的显式基元。一个学习任务首先声明一组持久指称对象和一个同一性准则;实现值$u$表示选中的指称对象。监督学习是主要的形式化特例,其语义对象是一族以单元为条件的响应律,同质性是这些响应律重合的特例;仅基于样本的条件无法判断世界是否同质,或观测到的响应律是否仅是异质族的边际分布。从数据中学习到的是一对$(T_\theta,R_\theta)$:一个生成上下文单元 token 的分词器,以及一个读取该 token 的共享响应律形式。结构化类将该形式视为 token 中的简单关系,线性预测器是其运行实例。token 是学习器侧的表征,任务侧的单元通过它影响预测;若学习器规范省略单元信息,则为单元不敏感;同质性仍是世界侧响应族的属性。当身份未明确时,世界侧的响应律混合以单元为条件的目标,学习器则将其共享形式与 token 组合。可信解析器可固定单元并提供查找 token;否则「单元溯因(unit abduction)」会从事实证据中生成同类型的 token。未关联的单行观测无法区分异质单元世界与同质池化世界;可信的同单元对可分离出受限的见证样本。形式化结果即关于该监督学习特例的研究。

英文摘要

Machine learning is usually formalized through samples, while the persistent individual to which multiple observed or possible events refer often remains implicit. We propose the \emph{unit} as an explicit primitive at the level of task semantics. A learning task first declares a population of persistent referents and a sameness criterion; the realized value $u$ denotes the selected referent. Supervised learning is the main formal specialization. Its semantic object is a family of unit-conditioned response laws. Homogeneity is the special case in which those laws coincide; a sample-only conditional is silent as to whether the world is homogeneous or the observed law is only the marginal of a heterogeneous family. What is learned from data is a pair $(T_ϕ,R_θ)$: a tokenizer that produces a contextual unit token and one shared response-law form that reads it. The structured class takes that form to be a simple relation in the token; a linear predictor is the running instance. The token is the learner-side representation through which the task-side unit affects prediction, while a learner specification that omits unit information is unit-insensitive; homogeneity remains a property of the world-side response family. When identity is unresolved, the world-side law mixes unit-conditioned targets, while the learner composes its shared form with a token. A trusted resolver may fix the unit and supply a lookup token; otherwise \emph{unit abduction} forms a token of the same type from factual evidence. Unlinked single-row observations can fail to distinguish a heterogeneous unit world from a homogeneous pooled world; trusted same-unit pairs separate a restricted witness. The formal results concern this supervised specialization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑