arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IKS-Instruct:一个用于教授语言模型印度知识体系的包含24000个示例的多语言数据集

IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

Shwetha Singaravelu, Gayathri Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani

arXiv 2607.23322首次发表:更新:

发表机构

Intrinsic Lab, Centre for Sensors, Instrumentation and Cyber-Physical System Engineering (Centre for SeNSE), Indian Institute of Technology Delhi; RSL Quantum, FITT, IIT Delhi(印度理工学院德里分校传感器、仪器仪表和网络物理系统工程中心(传感中心)本征实验室; 印度理工学院德里分校 FITT 学院 RSL 量子实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有指令数据集问题,提出IKS-Instruct多语言数据集教授语言模型印度知识体系,源于六种源类型,经多评判评估框架评估质量,展示了该数据集在特定模型微调上的效果及模型质量与数据整理的关系。

AI 中文摘要

指令微调已成为使大型语言模型遵循人类意图的标准方法,但现有指令数据集以英语常识任务为主,缺乏对专业教学领域的覆盖。本文提出IKS-Instruct,一个包含24795个指令-响应对的数据集,用于教授语言模型提供基于印度知识体系(IKS)的教育内容。该数据集涵盖七种语言,包含41种吠陀口头和数学传统的教学技术,与6至12年级的中央中等教育委员会(CBSE)课程一致。这些对来自六种源类型。通过多评判评估框架评估质量,在统一的五评判外部小组下,紧凑7B模型的最强IKS-Instruct微调达到中位数评判分数6.39,在特定IKS维度上,未进行IKS微调的基础模型分数接近零。模型质量并非随数据整理单调增加。

英文摘要

Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical domains. This paper presents IKS-Instruct, a dataset of 24,795 instruction-response pairs for teaching language models to deliver educational content grounded in Indian Knowledge Systems (IKS). The dataset spans seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam), covers 41 pedagogical techniques from the Vedic oral and mathematical traditions, and is aligned with the Central Board of Secondary Education (CBSE) curriculum for classes 6 through 12. The pairs are derived from six source types: classical text corpora (Bhagavad Gita, Thirukkural, Sangam literature, Vedic texts), curriculum-aligned pedagogical templates, Vedic mathematical sutra demonstrations, bilingual instruction pairs, technique-grounded multi-turn dialogues, and cross-tradition comparative analyses. Quality is assessed through a multi-judge evaluation framework in which independent language models score responses on 12 dimensions including technique fidelity, pedagogical quality, factual accuracy, and IKS cultural depth. Under a uniform five-judge external panel (median aggregation over 1,201 stratified items), the strongest IKS-Instruct fine-tune of a compact 7B model reaches a median judge score of 6.39, within 0.15 of a strong general-purpose reference model (Nemotron-Nano at 6.54) at a fraction of its deployment cost, while the base model without IKS fine-tuning scores near zero on the IKS-specific dimensions. Model quality does not increase monotonically with data curation, a result we report together with the corresponding data-quality gains.

Comments32 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑