发表机构
Carnegie Mellon University; Pratt Institute; University of Pennsylvania; University of Pittsburgh(卡内基梅隆大学; 普拉特学院; 宾夕法尼亚大学; 匹兹堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发布RegDivergence-101基准,提出跨司法管辖区监管分歧检测任务,对比四类方法性能,发现未提及类的检测特性,为大规模检测指明架构方向。
AI 中文摘要
为同时向美国和欧盟开发药物的制药申办方必须协调美国食品药品监督管理局(FDA)与欧洲药品管理局(EMA)各自发布的指南。若两家机构要求实质相同,申办方可仅提交一次申请;若要求存在分歧,单一试验设计可能被其中一个地区驳回;若某一机构对某一事项未作规定而另一机构有相关监管要求,申办方必须推断自身义务。目前,这种协调工作由监管事务专家手动完成。本文提出跨司法管辖区监管分歧检测任务:给定同一主题下的一项FDA要求和一项EMA要求,将二者关系分类为一致(AGREE)、分歧(DIVERGE)或未提及(SILENT)。未提及类具有内在方向性(分为SILENT_FDA和SILENT_EMA),本文记录每对的方向,并报告各方向的F1值及合并标签。本文发布RegDivergence-101,这是一个包含101对、经专家验证的试点评估基准(标签基于三项同行评审的FDA/EMA比较研究及FDA/EMA/ICH原始指南文本;双标注的标注者间kappa值为0.85),并系统表征了四类基线方法的性能层级:词汇启发式方法(宏F1值为0.511,95%置信区间为[0.411-0.605])、自然语言推理(NLI)交叉编码器(0.233)、义务级图检索增强生成(Graph-RAG,0.663,[0.570-0.747]),以及扁平大语言模型评判器Claude Haiku(0.830,[0.747-0.908])。在试点规模(n=101)下得出三项方向性观察结果:未提及类在语义上可检测,但仅依赖蕴含关系的模型无法识别;对级义务图性能优于词汇方法,但落后于扁平大语言模型上下文(置信区间部分重叠);语料库级图构建是大规模未提及检测的架构目标。RegDivergence-101是一份试点发布,确立了任务形式和基线层级,第7节描述了四个未涵盖的监管领域及扩展路线图。
英文摘要
Pharmaceutical sponsors developing a drug for both the United States and the European Union must reconcile guidance issued independently by the FDA and the EMA. Where the two agencies require substantively the same thing, a sponsor can file once; where they diverge, a single trial design risks rejection in one region; where one agency is silent on a point the other regulates, the sponsor must infer obligations. Today this reconciliation is performed manually by regulatory-affairs experts. We introduce cross-jurisdiction regulatory divergence detection: given an FDA requirement and an EMA requirement on the same topic, classify their relationship as AGREE, DIVERGE, or SILENT. SILENT is inherently directional (SILENT_FDA vs. SILENT_EMA); we record direction per pair and report per-direction F1 alongside the collapsed label. We release RegDivergence-101, a 101-pair expert-grounded pilot evaluation benchmark (labels grounded in three peer-reviewed FDA/EMA comparison studies and primary FDA/EMA/ICH guidance text; dual-annotation inter-annotator kappa = 0.85), and systematically characterise a four-method baseline hierarchy: lexical heuristic (0.511 macro-F1, 95% CI [0.411-0.605]), NLI cross-encoder (0.233), obligation-level Graph-RAG (0.663 [0.570-0.747]), and flat LLM judge / Claude Haiku (0.830 [0.747-0.908]). Three directional observations emerge at pilot scale (n = 101): SILENT is semantically detectable but invisible to entailment-only formulations; pair-level obligation graphs improve over lexical methods but trail flat-LLM context (CIs partially overlapping); and corpus-level graph construction is the indicated architectural target for large-scale silent-detection. RegDivergence-101 is a pilot release establishing the task formulation and baseline hierarchy; four unrepresented regulatory domains and an expansion roadmap are described in Section 7.