发表机构
Duale Hochschule Baden-Württemberg Ravensburg(巴登-符腾堡双元制应用技术大学 Ravensburg)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对德语文学文本在小语言模型相关应用的空白,提出Tiny_schiller语料库,经特定处理,可通过单文件和一行代码让小语言模型接触德语文学文本,填补了相关研究和应用的空白。
AI 中文摘要
Tiny_schiller填补了德语文学文本在小语言模型原型设计、微调、教育和研究方面的空白,提供了一个单文件的、可即插即用的对应于Karpathy的tiny_shakespeare的语料库。现有的德语文学语料库更大、更丰富,但在运行一行训练或微调代码之前需要解析器工程。tiny_schiller是一个2.07兆字节的单文件,包含十一部公共领域的席勒戏剧,源自DraCor的GerDraCor导出(CC0)并通过确定性解析器工程处理。通过从单个HuggingFace调用加载字符级、GPT-2字节对编码和cl100k_base令牌化分割、指令格式的对话完成分割以及89个每个字符的角色分割。一个小语言模型用一行代码就能接触到德语文学文本。
英文摘要
tiny_schiller closes the small-language-model prototyping, fine-tuning, education, and research gap for German literary text, providing a single-file, drop-in counterpart to Karpathy's tiny_shakespeare. The available German literary corpora are larger and richer, but require parser engineering before a single line of training or fine-tuning code can run. tiny_schiller is a 2.07-megabyte single file of eleven public-domain Schiller dramas, sourced from DraCor's GerDraCor export (CC0) and processed by deterministic parser engineering. Character-level, GPT-2 byte-pair encoding, and cl100k_base tokenization splits, an instruction-formatted dialogue-completion split, and 89 per-character persona splits load from a single HuggingFace call. A small language model literally reaches German literary text in one line of code.
Comments5 pages. Dataset: https://huggingface.co/datasets/mrkschtr/tiny_schiller ; Code: https://github.com/schutera/tiny_schiller