An individual successfully trained a 210 million parameter text-to-image diffusion transformer model from scratch using a single GPU over 3.5 days. The model, named TinyDiT, was trained on 4.2 million curated images and utilized a rectified flow approach with a FLUX.2 VAE and a frozen flan-t5-base for text encoding. Key factors for success included high-quality image-caption pairs, aspect-ratio bucketing, and the use of `torch.compile` for training acceleration. AI
IMPACT Demonstrates feasibility of training advanced diffusion models on consumer-grade hardware, potentially lowering barriers for independent AI research.
RANK_REASON The item describes the training of a novel diffusion transformer model from scratch by an individual, detailing the process and findings. [lever_c_demoted from research: ic=1 ai=1.0]
- flan-t5-base
- FLUX.2 VAE
- Hugging Face
- RTX PRO 6000
- Stable Diffusion
- text-to-image diffusion transformer
- torch.compile
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →