AI 中文总结
DualGuard提出双模式质量控制框架,通过选择性保留、过滤及历史行为对比实现逻辑保持的数据增强,在七个下游任务中五项最优并全面超越BERT基线。
AI 中文摘要
大型语言模型为大规模生成逻辑推理的增强数据提供了一种实用方法,但更大的生成量并不能保证语义、标签或逻辑上的可靠性。现有工作通过生成约束、候选验证、过滤和基于反馈的修订来提高生成质量;然而,一旦获得质量判断,决定候选应被保留、过滤还是修复仍然是一个重要的控制问题。我们提出DualGuard,一个用于逻辑保持数据增强的双模式质量控制框架。第一种模式使用当前实例和候选批次进行选择性保留、过滤、归因和定向反馈。第二种模式在逐样本诊断和归因的基础上,累积跨实例的增强操作执行记录,将新执行与每个操作自身的历史行为进行比较,并支持回顾性异常检查、定向回滚和有界修复。两种模式共享语义验证,并在有可靠逻辑形式时额外使用符号验证。在Two-Stage Transfer设置下的七个下游任务中,DualGuard在五个任务上取得了最高准确率,并在所有七个任务上优于无增强的BERT基线。受控消融进一步显示了Memory、Z3和历史感知异常控制的互补作用。
英文摘要
Large language models provide a practical way to generate augmented data for logical reasoning at scale, but a larger generation volume does not guarantee semantic, label, or logical reliability. Existing work has improved generation quality through generation constraints, candidate validation, filtering, and feedback-based revision; however, once a quality judgment is available, deciding whether a candidate should be retained, filtered, or repaired remains an important control problem. We propose DualGuard, a dual-mode quality-control framework for logic-preserving data augmentation. The first mode uses the current instance and candidate batch for selective retention, filtering, attribution, and targeted feedback. The second mode accumulates cross-instance execution records of augmentation actions on top of per-sample diagnosis and attribution, compares new executions against each action's own historical behavior, and supports retrospective anomaly inspection, targeted rollback, and bounded repair. Both modes share semantic verification and additionally use symbolic verification when a reliable logical form is available. Across seven downstream tasks in the Two-Stage Transfer setting, DualGuard achieves the highest Accuracy on five tasks and outperforms the no-augmentation BERT baseline on all seven. Controlled ablations further show complementary roles for Memory, Z3, and history-aware anomaly control.