arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

辅导大型语言模型实现领域自适应、精确与安全

Tutoring Large Language Models to be Domain-adaptive, Precise and Safe

Somnath Banerjee

arXiv 2609.23071首次发表:更新:

发表机构

Indian Institute of Technology, Kharagpur(印度理工学院卡拉格普尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本论文提出“负责任智能”框架,通过主动学习与图知识减少幻觉、解码时对齐阻止有害生成及语言特定引导确保文化安全,为构建情境化、伦理化且文化适应的下一代AI提供蓝图。

AI 中文摘要

本论文提出了一个“负责任智能”框架,以应对人工智能在安全、伦理和文化敏感性方面的关键挑战。该框架推进了三个核心领域:首先,通过主动学习和基于图的知识来改进专业领域的领域自适应,以减少幻觉。其次,通过一种新颖的解码时对齐机制增强伦理严谨性,该机制能够实时主动阻止有害文本的生成。最后,通过语言特定的引导确保文化和多语言安全,尊重多样化的语言和社会规范。最终,这项工作为构建具备情境知识、伦理健全和文化适应性的下一代人工智能提供了蓝图。

英文摘要

This thesis proposes a framework for "responsible intelligence" to address AI's critical challenges in safety, ethics, and cultural sensitivity. It advances three core areas: First, it improves domain adaptation in specialized fields using active learning and graph-based knowledge to reduce hallucinations. Second, it enhances ethical rigor via a novel decoding-time alignment mechanism that proactively blocks harmful text generation in real-time. Finally, it ensures cultural and multilingual safety through language-specific steering that respects diverse linguistic and social norms. Ultimately, this work provides a blueprint for building next-generation AI that is contextually knowledgeable, ethically sound, and culturally adaptable.

CommentsThis is a preprint of a PhD thesis submitted to IIT KGP. The final, official version of record is available through the university's institutional repository

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑