ATIBA:面向研究论文的基于依据的完整性与质量检查
ATIBA: Grounded Integrity and Quality Checking for Research Papers
- Bilkent University(比尔肯特大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ATIBA 是一款可对手稿执行五项基于依据的完整性与质量检查的工具,经 13 名非作者用户评估,其六项调查项一致性均值达 85%,初步验证了其感知有用性。
AI中文摘要:
检查手稿的参考文献完整性、其对目标 venues 特定投稿规则的合规性,以及对社区报告标准的遵守情况是一项手动、重复且因 venues 而异的工作,因此实际操作中往往不一致或被跳过。我们提出 ATIBA,这一工具可对手稿执行五项基于依据的完整性与质量检查:参考文献完整性检查,即对照书目来源验证每篇引用文献,并标记已撤回或无法找到的参考文献;venue/赛道合规性检查,即直接从 venue 的征稿启事页面提取投稿标准,并据此评估手稿,每项判定均锚定于该页面的逐字引用;ACM SIGSOFT 经验标准合规性检查,附带幻觉防御机制,会丢弃其无法在手稿中逐字找到的任何证据引用;由 Azure OpenAI 提供支持的 GPT-5.4 驱动的多模式 AI 评审(venue 特定、正式及锚定页面的标注);以及参考文献建议功能,即提出手稿的候选参考文献,并在向用户展示前对照书目来源验证每篇文献。所有五项检查均遵循同一原则:仅信任大语言模型(LLM)进行判定,绝不信任其生成判定所依据的证据。我们通过一项由 13 名非作者参与者参与的 moderated 用户研究对 ATIBA 进行评估。六项调查项的一致性介于 69% 至 92% 之间,均值为 85%,为所评估工作流程的积极感知有用性提供了初步证据。这些发现确立了感知有用性;客观准确性仍有待测量。
英文摘要:
Checking a manuscript's reference integrity, its compliance with a target venue's specific submission rules, and its adherence to community reporting standards is manual, repetitive, and different for every venue so in practice it is done inconsistently or skipped. We present ATIBA, a tool that runs five grounded integrity and quality checks on a manuscript: a reference-integrity check that verifies each citation against bibliographic sources and flags retracted or unfindable references; a venue/track compliance check that derives submission criteria directly from a venue's own call-for-papers page and evaluates the manuscript against them, each verdict anchored to a verbatim quote from that page; an empirical-standards compliance check against the ACM SIGSOFT Empirical Standards, with a hallucination defence that discards any evidence quote it cannot locate verbatim in the manuscript; a multi-mode AI review (venue-specific, formal, and page-anchored annotation) powered by GPT-5.4 through Azure OpenAI; and a citation-suggestion feature that proposes candidate references for a manuscript and verifies each against bibliographic sources before it is shown to the user. All five checks are designed around the same principle: an LLM is only trusted to judge, never to invent the evidence it judges against. We evaluated ATIBA through a moderated user study with 13 non-author participants. Agreement across the six survey items ranged from 69% to 92%, with a mean of 85%, providing initial evidence of positive perceived usefulness across the evaluated workflows. These findings establish perceived usefulness; objective accuracy remains to be measured.