5090
PulseAugur coverage of 5090 — every cluster mentioning 5090 across labs, papers, and developer communities, ranked by signal.
- 2026-08-27 product_launch The NVIDIA RTX 5090 graphics card has been officially priced at $5090. source
11 day(s) with sentiment data
-
NInfer fork boosts LLM context to 555k with 4-bit KV cache
A fork of the NInfer project has been developed, introducing significant improvements to context length and memory management for large language models. This fork features a custom 4-bit KV cache that reduces VRAM usage…
-
DGX Spark and 5090 prices rise due to corporate demand
The DGX Spark, a high-end computing component, is experiencing a price increase, mirroring a similar trend observed with the 5090. This surge in cost is attributed to significant demand from large corporations, leading …
-
NInfer boosts local LLM performance to 220 tokens/sec on 5090 GPU
A user on Reddit's r/LocalLLaMA subreddit shared their positive experience using NInfer with a 5090 GPU to run a 27 billion parameter model. They reported achieving significantly higher throughput, with speeds averaging…
-
NVIDIA RTX 5090 priced at $5090, sparking user concern
The NVIDIA RTX 5090 graphics card has been officially priced at $5090, a significant increase that has led some users to reconsider their hardware choices. One user noted that for the price of a 5090, they could instead…
-
Gigabyte Aorus Master 16 gaming laptop with OLED and RTX 5090 drops $800 to $3,599
A Gigabyte Aorus Master 16 gaming laptop, featuring an OLED display, an Intel Arrow Lake CPU, and an RTX 5090 mobile GPU, is available for $3,599, which is $800 off its original price. This high-performance laptop is eq…
-
DFlash2 shows speed gains but increased memory use for Qwen3.8 27B
A user tested DFlash2, a new method for accelerating large language model inference, with the Qwen3.8 27B model on an RTX 5090 GPU. While DFlash2 showed speed improvements, particularly for code generation, reaching up …
-
Genie world model runs 720p at 16 FPS on single consumer GPU
A new "Genie" world model has been demonstrated, capable of running at 720p resolution and 16 frames per second on a single consumer-grade GPU. This model utilizes 19GB of VRAM, making it accessible for users with high-…
-
Qwen3.8-27B model hits 880 tok/s on single RTX 5090 with NInfer engine
A user has achieved impressive performance with the Qwen3.8-27B model on a single RTX 5090 GPU, reaching 880 tokens/second with 4-bit NVFP4 quantization and a full 262k context. This speed was attained using the NInfer …
-
Minimax H3 on DGX Spark achieves efficient video generation, user seeks alternatives
A user on Reddit shared their experience running the Minimax H3 model on a DGX Spark, achieving impressive speed and cost-efficiency for video generation. They reported generating a 5-second clip in 203 seconds with opt…
-
GPU performance projections: 3x faster than 5090 by 2030?
A discussion on Reddit's r/LocalLLaMA forum speculates about future GPU performance increases. Based on an average generational jump of 50% since 2009, participants project that GPUs could be three times faster than a 5…
-
AI generates full Star Trek episode locally in one day
A user has created a full episode of "Star Trek: The Next Generation" locally in a single day using the MiniMax H3 model on a 5090 graphics card. This project generated native dialogue and audio without relying on separ…
-
Muse Glimmer 30B model hits 253 t/s on RTX 5090 with optimization
A user on Reddit's r/LocalLLaMA subreddit shared impressive performance benchmarks for the Muse Glimmer 30B model, achieving 253 tokens per second on an RTX 5090 GPU. This speed was attained using a specific quantizatio…
-
Glimmer LLM achieves 233.4 tps on 5090 GPU, reaches 256k context
A new model called Glimmer has demonstrated impressive performance, achieving 233.4 tps on a 5090 GPU with Dflash. Users are reporting that Glimmer can easily reach a 256k context window on 24GB of VRAM, a feat not easi…
-
H3 Turbo LoRA achieves fast 0.5MP video generation on RTX 5090
A Reddit user shared their experience using the H3 Turbo LoRA for Stable Diffusion, achieving fast generation times of 5 seconds of 0.5 megapixel video in 29 seconds on an RTX 5090 GPU. The user provided tips and settin…
-
User creates short film with MiniMax H3 R2V on local hardware
A user has created a short film using the MiniMax H3 R2V model locally on a 5090 graphics card. The process involved generating reusable references for characters, objects, and environments, then using H3 R2V to produce…
-
RTX 5090 power limit reduction yields minimal inference performance loss
A user on r/LocalLLaMA shared findings on reducing the power limit of an NVIDIA RTX 5090 graphics card for AI inference. By lowering the power limit to 480W, the card experienced only a negligible performance decrease o…
-
User seeks advice on dual PSU setup for high-power GPU workstation
A user on Reddit's r/LocalLLaMA forum is seeking advice on configuring a dual power supply unit (PSU) setup for a high-end workstation. The user plans to install multiple powerful GPUs, such as RTX Pro 6000s and 5090s, …
-
Krea2 user seeks VRAM optimization tips for RTX 5090 setup
A user on Reddit is seeking advice for optimizing their setup of Krea2, a diffusion model, on an RTX 5090 graphics card. They are experiencing VRAM limitations, requiring the text encoder to be reloaded from disk for ea…
-
AI Video Generation GPU Choice: RTX 5080 vs 3090 vs Pro 4000
A user on Reddit is seeking recommendations for a GPU to handle AI video generation, excluding the 4090 and 5090 models. They are considering the RTX 5080 16GB, RTX 3090 24GB, and the RTX PRO 4000 Blackwell 24GB, and ar…
-
User escalates local LLM hardware, finds initial setup sufficient
A user shared their experience of escalating hardware purchases for running large language models locally, starting with a single RTX 5090 for 27B models and fine-tuning. This led to acquiring two RTX 6000 Pros in antic…