arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01046cs.CLcs.LG

句子特异性评分用于协作技术文档:一项领域迁移研究

Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study

Rocker D'Antonio, Thomas Benton Townsend, Dimitrios Michael Manias

首次发表
浏览论文内容

中文总结 AI 辅助

本研究审计技术文档句子特异性评分工具,发现SpeciTeller在Gemma和GPT-OSS-120B集合中提升方向有效选择率,表明评分价值依赖预测器与候选集。

中文摘要 AI 辅助

协作依赖于共享上下文,而技术文档是上下文在人类和AI队友之间持续存在的一种方式。特异性,即语言中表达的细节数量和精确度,塑造了文档捕获哪些信息以及信息被传达的精确程度。本研究对技术文档上的句子特异性评分工具进行了审计,并测试了仅在生成后应用的评分是否有助于在固定的LLM生成的修订版本中进行选择。在维基百科和三个技术文档语料库中,固定的通用领域预测器SpeciTeller和Ko等人目标适应预测器的固定发布后作者-仓库实现产生了不同的语料库排序,同句排名一致性从-0.066到0.510。严格的过滤和令牌长度调整改变了这些模式,但没有调和它们。在Gemma集合中,SpeciTeller排名将方向有效选择从71.7%提高到83.3%(+11.7个百分点;95%源案例自举区间+1.7至+21.7);在GPT-OSS-120B集合中,SpeciTeller排名将方向有效选择从51.7%提高到56.7%(+5.0个百分点;95%源案例自举区间-6.7至+16.7),并且每个主要的单评分GPT-OSS-120B区间都包含零。这些发现将评分解释和决策价值与预测器和候选集联系起来。

英文摘要

Collaboration depends on shared context, and technical documentation is one way that context persists across people and AI teammates. Specificity, the amount and exactness of detail expressed in language, shapes what information documentation captures and how precisely that information is communicated. This work audits sentence-specificity scoring artifacts on technical documentation and tests whether scores applied only after generation help choose among fixed LLM-generated revisions. Across Wikipedia and three technical-documentation corpora, the fixed general-domain predictor SpeciTeller and the pinned post-publication author-repository implementation of Ko et al.'s target-adapted predictor produce different corpus orders and same-sentence rank agreement from -0.066 to 0.510. Strict filtering and token-length adjustment change these patterns without reconciling them. In the Gemma set, SpeciTeller ranking raises direction-valid selection from 71.7% to 83.3% (+11.7 points; 95% source-case bootstrap interval +1.7 to +21.7); in the GPT-OSS-120B set, SpeciTeller ranking raises direction-valid selection from 51.7% to 56.7% (+5.0 points; 95% source-case bootstrap interval -6.7 to +16.7), and every primary single-score GPT-OSS-120B interval includes zero. These findings tie score interpretation and decision value to the predictor and candidate set.

发表机构

  • Mississippi State University(密西西比州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑