PulseAugur
EN
LIVE 08:16:57

TINY_SCHILLER corpus simplifies German literary text for small language models

A new corpus named TINY_SCHILLER has been released, designed to facilitate the development and education of small language models using German literary text. This corpus is a single-file, drop-in replacement for the popular tiny_shakespeare dataset, offering eleven public-domain dramas by Friedrich Schiller. Processed for ease of use, TINY_SCHILLER supports various tokenization methods and persona splits, enabling researchers and students to work with German literary data in a single line of code. AI

IMPACT Enables easier research and education in German literary text for small language models.

RANK_REASON The item is an academic paper detailing a new dataset for language model research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

TINY_SCHILLER corpus simplifies German literary text for small language models

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Mark Schutera ·

    TINY_SCHILLER: A Drop-In German Drama Corpus for Small Language Models

    arXiv:2607.19992v1 Announce Type: cross Abstract: tiny_schiller closes the small-language-model prototyping, fine-tuning, education, and research gap for German literary text, providing a single-file, drop-in counterpart to Karpathy's tiny_shakespeare. The available German litera…