在块内思考:使用大语言模型生成符合法规的场景的RegulaRAG——以联合国第152号法规为例
Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152
浏览论文内容
中文总结 AI 辅助
针对LLMs难以结合冗长分层标准的问题,提出RegulaRAG流水线,经实验其在UN R152数据集上元分数最高且鲁棒性强,优于基线系统。
中文摘要 AI 辅助
生成符合法规的测试场景对于验证安全关键型汽车系统至关重要,但大语言模型(LLMs)难以将输出与冗长的分层标准相结合。我们提出RegulaRAG,这是一种检索增强生成(RAG)流水线,它结合了智能分块(SmartChunking),即通过图遍历对段落和表格进行参考感知增强,以及对这些增强单元的智能检索与重排序(Smart Retrieve & Rerank)。为测试该系统,我们在涵盖联合国第152号法规(AEBS)中所有场景的人工整理数据集上进行评估。本研究包括:(i)三步渐进式搜索,无需详尽网格搜索即可识别近最优检索参数;(ii)与五个基线RAG系统的直接对比;(iii)通过添加干扰内容扩展源语料库的鲁棒性压力测试。输出采用定制惩罚评分指标评估。在所有实验中,RegulaRAG的平均元分数最高(82.99),比次优系统高出43%(NoRAG:57.94),同时每查询处理14000至25000个标记,而以图为中心的基线系统高达500000个标记。它在法规源数量增加时仍保持稳定的高性能,而对比的RAG系统在质量和鲁棒性方面均大幅下降。
英文摘要
Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards. We present RegulaRAG, a Retrieval-Augmented Generation (RAG) pipeline that couples SmartChunking, reference-aware enrichment of paragraphs and tables via graph traversal, with Smart Retrieve & Rerank over these enriched units. To test our system, we evaluate on a manually curated dataset covering all scenarios in UN Regulation No. 152 (AEBS). Our study comprises: (i) a three-step progressive search that identifies near-optimal retrieval parameters without exhaustive grid search; (ii) head-to-head comparisons against five baseline RAG systems; and (iii) a robustness stress test that scales the source corpus with distractor content. Outputs are evaluated using a customized penalized scoring metric. Across all experiments, RegulaRAG achieves the highest average Meta-Score (82.99), outperforming the next-best system by 43% (NoRAG: 57.94), while operating at 14k-25k tokens per query versus up to 500k for graphcentric baselines. It maintains strong performance, remaining stable even as the number of regulatory sources grows, whereas competing RAG systems degrade sharply in both quality and robustness.
发表机构
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。