MI300X
PulseAugur coverage of MI300X — every cluster mentioning MI300X across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
AMD Instinct MI300X GPU detailed with custom Python management tools
This article details the process of inventorying and measuring an AMD Instinct MI300X GPU on the AMD Developer Cloud. The author developed a suite of Python tools, collectively named 'MCP', to manage and interact with t…
-
GPU rental prices show daily fluctuations across platforms · 8 sources tracked
GPU rental prices are fluctuating across various platforms like Vast.ai and RunPod, with daily updates tracking the cheapest options by VRAM. Prices for high-end GPUs such as the 192GB MI300X remain consistent at $2.39/…
-
LLM inference challenges: managing large models on limited GPU memory
Running large language models on consumer hardware presents challenges due to their significant memory requirements. Techniques like quantization, which reduces the precision of model weights, and sharding, which splits…
-
New frameworks optimize GPU kernels for deep learning and HPC · 3 sources tracked
Three new research papers introduce advanced frameworks for optimizing GPU kernels, crucial for deep learning and high-performance computing. HIERA focuses on workload-aware planning across different implementation spac…
-
Ruitong launches to measure LLM accuracy drift across hardware
A new initiative called Ruitong has been launched to address the issue of accuracy drift in large language models (LLMs) when migrated across different hardware. While latency and cost are well-understood, the impact on…
-
Proposal: Enforce AI Model Safety at the GPU Level
A proposal suggests enforcing AI model safety at the GPU level to mitigate risks, particularly from open-weight models. Current safety measures include internal model safeguards, input/output screening, and monitoring, …
-
Alibaba's Qwen3.8-27B model released; AI aids GPU porting; LLM infra detailed
Alibaba's Qwen team has released Qwen3.8-27B, a dense 27-billion parameter model that fits on a single GPU and supports a 1 million token context window, with Day-0 integration in vLLM. Concurrently, research is explori…
-
AMD MI355X performance boosted by community hackathon, rivals B200 on Kimi models · 6 sources tracked
AMD, in collaboration with GPU_MODE, has launched a $1.1 million kernel hackathon that has significantly improved the performance of its MI355X graphics card. The Readonflow Team's optimizations, focusing on MoE kernels…
-
AI's true bottleneck: Memory, not chips, limits LLM scalability
The AI industry is facing a critical bottleneck not in the availability of specialized chips like NVIDIA's H100 or AMD's MI300X, but in the memory systems that support them. While companies like Google, Microsoft, and M…
-
Unsloth adds AMD GPU support for faster local LLM training
Unsloth has released an update that significantly enhances support for AMD GPUs, enabling local LLM training and inference across various AMD hardware. This new version promises up to 2x faster performance and 70% less …
-
Inferra proposes GPU compute futures exchange to tackle fragmented market
The procurement of GPUs for AI development remains challenging due to fragmented access, uneven allocation of high-demand chips like H100s, and a lack of price transparency across providers. Existing solutions such as r…
-
Machine0.io launches persistent VMs with CLI control
Machine0.io has launched a new service offering persistent virtual machines (VMs) for developers and agents, accessible via a command-line interface (CLI). These VMs run NixOS or Ubuntu with pre-installed tools, providi…
-
Fine-tune LLMs on AMD MI300X using ROCm and QLoRA
This article details a practical workflow for fine-tuning large language models using AMD's ROCm platform, specifically on the MI300X hardware. It highlights how to overcome the dominance of NVIDIA's CUDA by leveraging …
-
TritonMoE kernel enables cross-platform MoE inference
Researchers have developed TritonMoE, a new inference kernel for Mixture-of-Experts (MoE) models written entirely in OpenAI's Triton language. This kernel achieves cross-platform compatibility, running on both NVIDIA an…
-
AMD's MI300X falls short in AI training due to software issues
A recent benchmark analysis reveals that AMD's MI300X, despite theoretical advantages in specifications and total cost of ownership, is not competitive with NVIDIA's H100 and H200 for AI training workloads. The primary …