A developer has created a new AI world model capable of generating playable characters from images, with the ability to switch prompts mid-generation. This model, named LocalAI World Model Part 2, utilizes a pure transformer architecture with a block causal mask and a diffusion forcing training method. It can run in real-time on consumer hardware like an RTX 5090 or M5 MacBook, achieving high frame rates by using a sliding window for context, similar to LLMs. The model also incorporates text cross-attention and has undergone extensive text-video pretraining, allowing it to respond to dynamic prompt changes such as altering environments or character appearances. AI
IMPACT This model could enable more interactive and dynamic AI-driven character creation for gaming and virtual environments.
RANK_REASON The item describes a novel AI model architecture and its capabilities, akin to a research demonstration. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →