arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

两种词汇,一种现象:人工智能在生育率下降证据综合中的元数据偏差

Two Vocabularies, One Phenomenon: Metadata Bias in AI Evidence Synthesis on Fertility Decline

Boyana Buyuklieva, Veronica-Nicolle Hera

arXiv 2607.09409首次发表:更新:

AI 中文总结

研究人工智能在生育率下降证据综合中的元数据偏差,通过 OpenAlex 对同一现象的临床和社会篮子查询,比较元数据等方面,发现社会框架机器可读性低,开放获取率相当,偏差存在于索引深度,影响语言模型提取假设和主张。

AI 中文摘要

生育率下降是未来十年关键政策问题之一,政策制定者对其的认知 increasingly 受人工智能合成证据库影响。但用此类工具询问生育相关问题时,答案取决于用词。同一现象,临床表述(如不孕、试管婴儿)或社会表述(如无子女、生育意愿)在目录中的完整性差异巨大。虽数据库对社会科学、书籍和灰色文献索引不足早有定论,但本文新之处在于固定主题,探究元数据差距是否对生育决定因素这一争议问题构成隐藏政策过滤器。我们用 OpenAlex 对同一现象进行两个查询:临床篮子(不孕、亚生育、辅助生殖技术、试管婴儿、生育力;n = 101,645)和社会篮子(无子女、社会不孕、生育意愿、生殖决策;n = 3,646)。比较它们在元数据完整性、开放获取、输出类型和机构来源方面的情况。社会框架的机器可读性始终较低:输出偏向书籍和论文,作者来自大学而非医疗机构。开放获取率基本相等(43.1%对41.3%),差距在于索引深度而非付费墙,这表明简单的开放获取指令无法解决问题。在这种混合文献上,甚至在任何覆盖偏差出现之前,当被要求提取假设和主张时,语言模型工具错过的比捕捉到的更多;本文记录的偏差使本就不完美的提取阶段更加复杂。

英文摘要

Declining fertility is one of the defining policy questions of the next decade, and increasingly, what policymakers know about it is shaped by AI-synthesising the evidence base. But ask such a tool about reproduction and the answer depends on the word you use. The same phenomenon, framed clinically (e.g, infertility, IVF) or socially (e.g., childlessness, fertility intentions), is catalogued with radically different completeness. And the catalogue as much, if not more than the underlying scholarship, is what AI synthesis begins with (Bolaños et al., 2024). That databases under-index the social sciences, books, and grey literature is well established (Visser et al., 2021). What is new here is holding the topic fixed and asking whether metadata gaps act as a hidden policy filter on a single contested issue: the determinants of (in)fertility. We use two OpenAlex queries on the same phenomenon: a clinical basket (infertility, subfertility, ART, IVF, fecundity; n=101,645) and a social basket (childlessness, social infertility, fertility intentions, reproductive decision-making; n=3,646). We compare them on metadata completeness, open access, output type, and institutional provenance. The social framing is consistently less machine-legible: output skewed to books and dissertations, authorship university -- rather than healthcare-based. Open access rates are essentially equal (43.1% vs 41.3%), so the gap is in indexing depth, not paywalls, suggesting simple OA mandates will not fix it. On this same mixed literature, even before any coverage bias enters the picture, LLM tools already miss more than they catch when asked to extract hypotheses and claims (Uprety et al., 2025); the bias documented here compounds an already-imperfect extraction stage.

Comments6 pages, 4 figures, Submission to Data for Policy 2026 conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑