Apple researchers have developed STARFlow2, a novel architecture that unifies multimodal generation by bridging language models and normalizing flows. This approach allows for continuous, single-pass, and purely causal generation of interleaved text and image sequences, preserving pretrained multimodal understanding and enabling high-fidelity image synthesis. The system utilizes a Pretzel architecture with residual skip connections and a unified latent space, allowing both text and visual outputs to enter the KV-cache directly, thereby improving generation efficiency. AI
IMPACT This research could lead to more integrated and efficient multimodal AI systems capable of generating coherent text and images.
RANK_REASON The cluster contains a research paper detailing a new multimodal generation model. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Apple Machine Learning Research →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →