SupraLabs has released a new dataset called reasoning-corpus-4K-5M-v1, containing 5 million samples designed to help train smaller language models in reasoning capabilities. The dataset includes user prompts, detailed thought traces, and AI model answers, all formatted in ChatML and within a 5k sequence length for efficient fine-tuning. This resource is available on Hugging Face and has already seen significant early adoption. AI
IMPACT Provides a specialized dataset to improve reasoning capabilities in smaller language models, potentially enabling more efficient and capable on-device AI.
RANK_REASON Release of a new dataset for training AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →