A new corpus named TINY_SCHILLER has been released, designed to facilitate the development and education of small language models using German literary text. This corpus is a single-file, drop-in replacement for the popular tiny_shakespeare dataset, offering eleven public-domain dramas by Friedrich Schiller. Processed for ease of use, TINY_SCHILLER supports various tokenization methods and persona splits, enabling researchers and students to work with German literary data in a single line of code. AI
IMPACT Enables easier research and education in German literary text for small language models.
RANK_REASON The item is an academic paper detailing a new dataset for language model research. [lever_c_demoted from research: ic=1 ai=1.0]
- DraCor
- Friedrich Schiller
- GerDraCor
- German
- GPT-2
- Hugging Face
- Karpathy
- Mark Schutera
- small language model
- TINY_SCHILLER
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →