RTX 3090
PulseAugur coverage of RTX 3090 — every cluster mentioning RTX 3090 across labs, papers, and developer communities, ranked by signal.
- used by Qwen3.8-27B 90%
- used by CodeLlama 34B 90%
- instance of RTX 4090 70%
- used by Qwen3.6-27B 70%
- used by vLLM 70%
- competes with RTX 5090 70%
- used by Gemma 70%
- used by Qwen3.6 35B-A3B 70%
- competes with GeForce RTX 3060 70%
- used by Multi Token Prediction 70%
- used by Qwen-3.6 27B 70%
- used by r/LocalLLaMA 70%
12 day(s) with sentiment data
-
Local LLM Hardware: GPUs for Small Models, Unified Memory for Large
For running large language models locally, the hardware landscape has divided into distinct categories based on memory capacity and speed. Consumer GPUs like the RTX 5090 excel with smaller models fitting within 32 GB, …
-
Qwen 3.8-27B attempts Riemann hypothesis over 63 hours
A user ran the Qwen 3.8-27B model for 63 hours on an RTX 3090 to attempt to solve the Riemann hypothesis. While the model did not succeed, the experiment demonstrated its ability to organize information, correct its own…
-
Perplexity's local AI agent now available on Windows for RTX GPUs
Perplexity has launched its Portable Computer AI agent for Windows, enabling users to run multistep AI tasks locally on their PCs. This feature, previously exclusive to Linux, requires a GeForce RTX or RTX PRO GPU with …
-
New tools and research tackle GPU optimization for AI workloads
Several research papers and a new open-source tool address challenges in optimizing AI workloads on GPUs. COMPASS-ABS aims to reduce fragmentation in shared GPU clusters for deep learning training, improving resource ut…
-
Qwen3.8 27b model tested for game development capabilities
A user on Reddit has been experimenting with the Qwen3.8 27b language model, specifically testing its capabilities in game development. The user has built a game using the model, which took approximately five hours to c…
-
Developer upgrades to RTX 3090 for faster local LLM performance
A software developer has transitioned from a MacBook M5 Pro with 48 GB of RAM to a custom-built Linux machine featuring an RTX 3090 graphics card. This upgrade has significantly boosted their local large language model …
-
vLLM enables 144K context Qwen3.8 27B on RTX 3090
A user on Reddit's r/LocalLLaMA subreddit shared a method for running the Qwen3.8 27B model with a 144K context window on an RTX 3090 GPU using vLLM. The user detailed a process involving Ahead-of-Time (AOT) compilation…
-
Reddit user compiles GPU guide for local AI, focusing on GB/dollar and bandwidth
A Reddit user has compiled a guide comparing GPUs based on their GB per dollar and bandwidth, focusing on models frequently discussed in local AI communities. The script used to gather this data pulls from subreddits li…
-
Qwen 3.8 27B benchmarks show software bottlenecks limit performance on high-end GPUs
New benchmarks reveal that while Alibaba's Qwen 3.8 27B model shows promise, its performance is significantly hampered by software and inference engine bottlenecks, rather than VRAM capacity. Testing on high-end GPUs li…
-
Local AI assistant runs Qwen3.8-27B model on dual RTX 3090s
A user has successfully set up a local AI assistant named Jarvis using the Qwen3.8-27B model, which runs on two RTX 3090 GPUs. This setup allows for daily tasks such as processing Jira emails, generating compliance tabl…
-
User seeks advice on PC build for local LLM inference
A user is seeking advice on building a PC for local large language model (LLM) inference. They are debating between two hardware configurations: a more budget-friendly option using dual NVIDIA RTX 3060 GPUs with 12GB VR…
-
Used RTX 3090 cards offer potential for local AI, but require careful evaluation
Used RTX 3090 graphics cards can be a viable option for local AI workloads, but potential buyers should first assess their specific needs. Factors such as runtime builds, available VRAM, power consumption, and the condi…
-
RTX 3090's 24GB Memory Remains Crucial for Local AI in 2026
Used RTX 3090 graphics cards will continue to be relevant for local AI applications in 2026 due to their 24GB of CUDA memory. However, running dual GPUs, managing KV cache, and PCIe bandwidth limitations will be key con…
-
User tests MiniMax H3 settings on RTX 3090 for optimal image generation
A user on Reddit conducted 112 tests using the MiniMax H3 model on an RTX 3090 graphics card to identify optimal settings for image generation. The tests systematically varied model/LoRA combinations, resolution, and st…
-
Qwen models dominate local LLM downloads, surpassing Llama and Meta
As of September 2026, the landscape of locally runnable large language models has shifted significantly, with Chinese models like Qwen dominating downloads and usage on platforms such as Hugging Face, surpassing Meta's …
-
Local AI setup outperforms expensive SaaS customization
An individual found that customizing existing SaaS software was prohibitively expensive and imperfect, costing between €500-€1500. In response, they developed their own custom application using local AI tools, an RTX 30…
-
AMD GPU Choice for AI Workloads: R9700 vs W7800
A user is seeking advice on choosing between two AMD GPU configurations for their workstation, aiming to enhance local inference and PyTorch training capabilities. The options are two Radeon AI PRO R9700 32GB cards or t…
-
LLMs accelerate Persian chat anonymization with efficient NER training
Researchers have developed a method for efficient anonymization of Persian customer chats using LLM-labeled data. They compared three instruction-tuned LLMs—DeepSeek-V3-0324, GPT-OSS-120B, and Qwen3-235B-A22B-Instruct-2…
-
LLM Enthusiasts Question Lack of INT8 W8A8 Model Adoption Despite RTX 3090 Support
A discussion on Reddit explores why INT8 W8A8 models are not more prevalent among LLM enthusiasts, despite the RTX 3090 being a popular GPU with native INT8 tensor cores that could offer performance benefits. Users spec…
-
Qwen3.8-27B model achieves 2,000 prefill tokens/sec on RTX 3090
A user has achieved significant performance gains for the Qwen3.8-27B large language model on an RTX 3090 graphics card. By implementing a custom kernel that maintains near-fp32 quality at int8 precision, they boosted p…