From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
机构 * Yang Li Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences University of Chinese Academy of Sciences(中国科学院计算技术研究所、中国科学院自动化研究所、中国科学院大学) ; Qiang Sheng Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所、中国科学院自动化研究所) ; Yehan Yang Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所、中国科学院自动化研究所) ; Xueyao Zhang The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Juan Cao Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences University of Chinese Academy of Sciences(中国科学院计算技术研究所、中国科学院自动化研究所、中国科学院大学)
专题命中 安全训练 :alignment(abstract);DPO(abstract);safety(abstract);harmlessness(abstract)
Comments NeurIPS 2025 Accepted Paper