diffusers
PulseAugur coverage of diffusers — every cluster mentioning diffusers across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Prism model enables 2K video-audio generation with dynamic sparse attention
A new model named Prism has been released, designed for native joint video-audio generation at resolutions up to 2K. Developed by researchers from Fudan University and Tencent Hunyuan, Prism utilizes a dynamic sparse at…
-
Hugging Face highlights AI inference speedups and new techniques · 3 sources tracked
Hugging Face is highlighting several advancements in AI inference and speed. The platform is showcasing "nunchaku," a 4-bit diffusion inference technique integrated into its diffusers library. Additionally, Baseten is n…
-
OpenVDN releases VDN-Minimax-H3 for faster video generation
OpenVDN has released VDN-Minimax-H3 (VDN-H3), an open-source video generation model that utilizes a hybrid-attention architecture for faster inference. The model can generate a 14.4-second video clip in 11.23 seconds on…
-
Viggle-Animate enables fast video character replacement without pose estimation
Viggle-Animate is a new AI model that enables character replacement in videos by using a single repainted frame. This model, a finetune of MiniMax-H3's ref2va transformer, bypasses the need for intermediate representati…
-
LLaDA-Image: Open-source model for unified image generation and editing released
inclusionAI has released LLaDA-Image, an open-source unified model for image generation and editing. The model family includes a 50-step Base model for high-quality results and a 4-step distilled Turbo model for faster …
-
Alibaba-pai releases MiniMax-H3-Acc-LoRAs for efficient video generation
Alibaba-pai has released MiniMax-H3-Acc-LoRAs, a model designed for efficient video generation. This model utilizes Parallel Decoding Distillation (PDD) to achieve fast video creation in a minimal number of inference st…
-
Hugging Face Spotlights Diverse AI Tools and Projects
Hugging Face is highlighting various AI projects and tools through its blog. Recent posts feature Falcon Perception from TII UAE, DeepInfra's role as an inference provider, and Hcompany's AI browser partner, HoloTab. Ad…
-
MiniMax AI releases open-weight music model and tops video editing benchmark
MiniMax AI has released Minimax Music 3, an open-weight music generation model that utilizes an 8B LLM and a 2.7B Diffusion Transformer. This model is capable of producing full songs from text prompts and lyrics, and is…
-
MiniMax H3 model sees major speed boost with Sol Engine
MiniMax AI has announced significant speed improvements for its MiniMax H3 model using the Sol Engine. This agent-native Sol Video Inference Engine achieved a 3.95x speedup compared to Diffusers and a 2.80x speedup over…
-
Local AI Updates: llama.cpp, PyTorch, Kimi-K3, and NVIDIA NeMo Speech 3.0
Recent updates in the local AI and open-source model space include performance enhancements for llama.cpp with CUDA fusion, addressing critical quantization bugs in PyTorch for AMD GPUs, and the trending Moonshot AI Kim…
-
MiniMax Music 3 model generates full songs, leveraging Qwen3-8B LLM
MiniMax Music 3, a new open-weight model capable of generating complete songs up to five minutes long, has been released. This model utilizes a hybrid approach, combining an 8B Global LLM initialized from Qwen3-8B for l…
-
llama.cpp optimizes for Apple Silicon, Hugging Face boosts 4-bit diffusion inference
The latest release of llama.cpp, version b10299, introduces optimizations for Apple Silicon, enhancing performance on macOS and iOS devices using the Metal API. Additionally, Hugging Face has detailed its Nunchaku 4-bit…
-
Hugging Face highlights new inference providers and AI tools · 4 sources tracked
Hugging Face is highlighting several companies and projects that are enhancing its inference capabilities. DeepInfra has been featured as an inference provider, while Hcompany's HoloTab is introduced as an AI browser pa…
-
MiniMax AI releases open weights for H3 video model, igniting community innovation
MiniMax AI has released the open weights for its MiniMax-H3 video model, sparking rapid innovation within the open-source community. Within 48 hours of the release, users demonstrated the model's capabilities on unconve…
-
MiniMax H3 video model released with open weights, accessible via platforms
MiniMax AI has officially released its new omni-modal generative system, MiniMax H3, which can produce video with synchronized stereo audio up to 2K resolution and 15 seconds in duration. The model is now publicly avail…
-
New AI techniques enable local diffusion models and cost-efficient coding
A new 4-bit quantization technique called Nunchaku has been integrated into Hugging Face's Diffusers library, significantly reducing the VRAM needed for diffusion models without sacrificing quality. This makes powerful …
-
Hugging Face integrates Nunchaku for 4-bit diffusion inference
Hugging Face has introduced Nunchaku, a new method for 4-bit diffusion inference. This technique is integrated into the diffusers library, aiming to improve the efficiency of AI-generated content.
-
Microsoft releases Mage-Flow, a compact 4B image generation model
Microsoft has released Mage-Flow, a compact 4B-scale generative model designed for efficient text-to-image generation and instruction-based image editing. The model achieves competitive quality through a co-designed tok…
-
PEFT methods offer efficient fine-tuning for large language models
Parameter-Efficient Fine-Tuning (PEFT) offers a way to adapt large pre-trained models to new tasks by training only a small subset of parameters or adding lightweight components. This approach, distinct from full fine-t…
-
NVIDIA releases Nemotron 3.5 safety model and NeMo AutoModel for large-scale fine-tuning · 2 sources tracked
NVIDIA has released Nemotron 3.5, a multimodal safety model designed for global enterprises. This model offers customizable content safety solutions. Additionally, NVIDIA's NeMo AutoModel, in conjunction with Hugging Fa…