variational auto-encoder
PulseAugur coverage of variational auto-encoder — every cluster mentioning variational auto-encoder across labs, papers, and developer communities, ranked by signal.
- instance of Variational Autoencoders 90%
- instance of autoencoder 90%
- instance of Gaussian function 90%
- instance of Gotit.pub 70%
- used by alphaXiv 70%
- used by Diffusion Transformer 70%
- used by CatalyzeX 70%
- uses generative adversarial network 70%
- affiliated with generative adversarial network 70%
- instance of Latent diffusion model 70%
- used by Flux 70%
- used by DagsHub 60%
23 day(s) with sentiment data
-
LTX 2.5 users see major speedups with ConvRot video VAE
A user on Reddit's r/StableDiffusion subreddit shared a tip for improving generation times with LTX 2.5. By switching to a ConvRot video VAE, users can significantly reduce the time it takes to generate images. The post…
-
Flex-π model integrates 3D geometry and object semantics with RGB data
Researchers have developed Flex-$\pi$, a 6-billion parameter world-action model that integrates 3D geometry and object semantics alongside RGB data. This model leverages a pre-trained video-generation VAE to encode 3D p…
-
New Latent Bridge Matching framework synthesizes breast MRI from pre-contrast images
Researchers have developed a new framework called Latent Bridge Matching (LBM) for synthesizing contrast-enhanced breast DCE-MRI from pre-contrast images. This method utilizes a latent diffusion model (LDM) approach but…
-
New ELVAE model enhances uncertainty-aware generation in VAEs
Researchers have developed ELVAE, a new variational autoencoder that incorporates evidential learning to better distinguish between uncertainty in latent representations and variability around them. This approach allows…
-
New research proposes unsupervised disentanglement via functional orthogonality
A new research paper proposes a novel approach to unsupervised disentangled representation learning by framing latent concepts as factors influencing observations through locally orthogonal directions. This method, form…
-
UniSpace introduces unified visual representation for AI generation and editing
Researchers have developed UniSpace, a novel approach to visual representation that unifies understanding, generation, and editing tasks within a single model. By introducing "Patch Reparameterization," UniSpace modifie…
-
CuteTTS system enhances speech synthesis with continuous autoregressive modeling
Researchers have developed CuteTTS, a novel text-to-speech system designed for efficient and high-quality voice synthesis. This system utilizes continuous autoregressive modeling with variational auto-encoder latents an…
-
Latent-to-4D method enables reusable 3D scene generation from video latents
Researchers have developed Latent-to-4D, a novel method for generating dynamic 3D scenes from text or images. This approach bypasses the need to reconstruct RGB videos by directly aligning video diffusion latents with a…
-
MiniMax H3 open-weights video AI rivals Sora2, runs locally
MiniMax has released its H3 video generation model with open weights, enabling users to run it locally on consumer GPUs. The model supports text, image, video, and audio inputs, producing up to 15-second clips with ster…
-
Comfy adds Int8 ConvRot VAE support with Minimax H3 video VAE conversion
Kijai has introduced support for Int8 ConvRot VAEs within the Comfy framework. This update includes the conversion of the Minimax H3 video VAE, with the converted model available on Hugging Face.
-
Vorch-Omni framework unifies audio-visual generation tasks
Researchers have introduced Vorch-Omni, a unified framework designed for multi-task audio-visual synthesis. This system can handle a wide array of tasks, treating both video and audio signals as either inputs or outputs…
-
Flash-VAED framework accelerates video generation by 6x
Researchers have developed Flash-VAED, a framework designed to accelerate the VAE decoders used in latent diffusion models for video generation. This approach employs channel pruning and dominant operator optimization t…
-
New KVAE tokenizers aim to advance multimodal generative models
Researchers have introduced a new family of tokenizers called KVAE, designed for multimodal generative models. These tokenizers, including KVAE-Audio, KVAE-3D, and KVAE-2D, are specifically engineered for text-condition…
-
New autoregressive transformer model for single-cell gene expression generation
Researchers have developed a novel autoregressive transformer model designed for generating single-cell gene expression vectors. This model, which incorporates a learned quantized variational auto-encoder tokenizer, is …
-
SynAgent framework enables scalable cooperative humanoid manipulation
Researchers have introduced SynAgent, a novel framework designed to enhance cooperative humanoid manipulation capabilities. This system addresses data scarcity and coordination complexities by transferring skills from s…
-
New AI watermarking techniques emerge amid security concerns · 3 sources tracked
Researchers have developed new methods for watermarking AI-generated images to ensure authenticity and prevent forgery. One approach, IRIS, binds watermarks to the image's visual semantics, making them resistant to tran…
-
Generative AI framework creates realistic maritime safety scenarios for autonomous ship testing
Researchers have developed a new generative AI framework designed to create realistic and diverse safety-critical encounter scenarios for the digital testing of autonomous maritime navigation systems. This framework con…
-
FlexComposer framework unifies video compositing with trajectory control
Researchers have introduced FlexComposer, a novel framework designed to unify video compositing tasks. This system allows for the seamless integration of both static images and dynamic footage into existing video sequen…
-
Multiple Android phones tested to run large Stable Diffusion models
An experimental setup is testing the ability to run large Stable Diffusion models across multiple Android smartphones, aiming to overcome memory limitations on single devices. This approach, while not intended to outper…
-
Kandinsky 5 open weights require user integration, not just closed issues
While several runtime issues in the Kandinsky 5 repository have been closed, this does not guarantee that the model will run locally. The open-weights nature of the model shifts the integration burden to the user, requi…