PulseAugur
EN
LIVE 22:04:34

User trains game music generator Localsong with 1.2B DiT architecture

A user has developed and trained a game music generator model called Localsong, which is based on the 1.2B DiT architecture. This model was trained over eight days using a single H100 cloud GPU and incorporates the VAE from Stable Audio 3. The creator aims for Localsong to produce a broader spectrum of instrumental music styles compared to existing models like Ace-Step, Minimax M3, and Stable Audio 3, and has made the model and a WebUI available on Hugging Face. AI

IMPACT This user-developed tool demonstrates the accessibility of training specialized generative models for niche applications like game music.

RANK_REASON User-developed tool release not from a major AI lab.

Read on r/StableDiffusion →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

User trains game music generator Localsong with 1.2B DiT architecture

COVERAGE [1]

  1. r/StableDiffusion TIER_2 English(EN) · /u/Amazing-You9339 ·

    I trained a game music generator

    <!-- SC_OFF --><div class="md"><p>I trained a instrumental game music generator. The 1.2B DiT was trained on 1 cloud H100 from scratch in 8 days; I used the VAE from Stable Audio 3.</p> <p><a href="https://huggingface.co/Localsong/Localsong">https://huggingface.co/Localsong/Local…