graphics processing unit
PulseAugur coverage of graphics processing unit — every cluster mentioning graphics processing unit across labs, papers, and developer communities, ranked by signal.
- instance of Blackwell 90%
- instance of Nvidia Rubin Gpu 90%
- instance of Nvidia L4 90%
- used by Nvidia L4 90%
- developed by Flashattention 90%
- instance of Nvidia A100 90%
- used by DagsHub 70%
- used by Gotit.pub 70%
- used by alphaXiv 70%
- used by CatalyzeX 70%
- used by ScienceCast 70%
- used by High Bandwidth Memory 70%
30 day(s) with sentiment data
-
Multi-LoRA Serving Latency Solved with Dependency-Aware Caching
Startups fine-tuning large language models for specific customer needs often face escalating infrastructure costs. A common solution is to use Multi-LoRA serving, which allows multiple fine-tuned adapters to run on a si…
-
Geekbench 7 overhauls benchmarks with AI, CUDA, and real-world testing
Primate Labs has released Geekbench 7, a significant update to its cross-platform benchmarking tool. The new version features more realistic workloads for CPUs and GPUs, including AI-specific tasks like real-time face t…
-
NVIDIA GPUs set for lunar deployment to power space exploration
NVIDIA is expanding its GPU deployment to the moon, with its Jetson chips slated for use in lunar rovers and orbiting satellites. This initiative aims to leverage edge AI and specialized hardware for future space explor…
-
Nvidia sends GPUs to Moon; AI chip startup Etched valued at $10.3B
Nvidia is participating in a lunar mission by sending graphics processing units (GPUs) to the Moon. Separately, AI chip startup Etched has achieved a valuation of $10.3 billion, overcoming skepticism from investors.
-
SPORD method enhances e-commerce supply chain planning with AI
Researchers have developed SPORD, a novel approach for supply chain planning that combines simulation and optimization to address challenges in e-commerce logistics. This method, implemented as JD.com's NetSim platform,…
-
NVIDIA urges AI model co-design to boost GPU utilization
NVIDIA has released a technical blog post highlighting a critical issue in AI model design: poor hardware utilization due to models not being optimized for GPU architecture. The post explains that concepts like 'arithme…
-
New ECRAM system accelerates edge continual learning, slashing energy use
Researchers have developed CLASP, a novel system designed to accelerate continual learning on edge devices by integrating in-memory computing (IMC) with a specialized ECRAM device. This approach addresses the significan…
-
New SubQuad pipeline improves adaptive immune repertoire analysis
Researchers have developed SubQuad, a new pipeline designed to overcome limitations in analyzing adaptive immune repertoires. This system addresses the computational cost of pairwise affinity evaluations and dataset imb…
-
New AI Framework Accelerates Quantum Optimization for Complex Problems
Researchers have developed DQAOA-GPT, a novel framework that combines a distributed quantum approximate optimization algorithm with GPT-based quantum circuit generation. This hybrid approach aims to solve complex combin…
-
New MoE technique boosts LLM efficiency with compute-communication overlap
Researchers have developed a novel method to improve the efficiency of Mixture-of-Experts (MoE) models, which are crucial for scaling large language models. Their approach focuses on overlapping expert computation with …
-
New research optimizes transformer attention with Mathematics of Arrays
A new research paper details a method for optimizing transformer attention inference using the Mathematics of Arrays (MoA). The paper presents four memory-efficient artifacts, including a single-query decode DNF that al…
-
New LO-FAR workflow offers cost-effective feature ranking for ad recommendations
Researchers have developed LO-FAR, a new workflow for ranking sparse features in industrial ad recommendation systems. This CPU-only method uses lightweight local estimators to rank features based on their predictive si…
-
Chinese telcos embrace TPUs for AI compute, focusing on token economics
Chinese telecom operators are increasingly investing in AI computing power, moving beyond traditional GPUs to explore other chip architectures like TPUs and ASICs. This shift is driven by a focus on cost-efficiency, par…
-
User laments slow local AI model performance and hardware strain
The user expresses frustration with the performance of local AI models, comparing them to slow torrent servers that strain computer hardware like GPUs. They highlight the slow processing speed of 4 tokens per second and…
-
AI agents to drive significant CPU demand, potentially reshaping market value
A recent report from CITIC Securities predicts that AI agents will significantly increase demand for central processing units (CPUs), potentially leading to a revaluation of their market worth. The report suggests that …
-
Elon Musk confirms Micron chip allocation to Tesla amid HKEX market pressure
Elon Musk announced that Micron Technology has allocated a significant quantity of memory chips to Tesla, securing favorable terms despite current market scarcity. This allocation comes as the Hong Kong Stock Exchange f…
-
Hong Kong AI, GPU stocks face unlock pressure; Chinese AI model nears Opus 4.8
A new report indicates that the Hong Kong stock market is preparing for a significant test in the latter half of 2026, with substantial market value from AI large models, GPUs, and high-end semiconductors set to be unlo…
-
Critical minerals control AI's future via GPU and data center supply chains
The AI industry's future is significantly influenced by critical mineral supply chains, with six key chokepoints identified. These points control essential components like GPUs, HBM chips, and data center cooling system…
-
AWQ outperforms GPTQ in 4-bit quantization for local LLMs, but GPU and kernels are key
A comparison of 4-bit quantization methods for local Large Language Models (LLMs) indicates that Activation Aware Quantization (AWQ) generally outperforms GPTQ. However, the study emphasizes that the actual performance …
-
SkewAdam optimizer slashes MoE training memory by 97%
A new optimizer called SkewAdam has been developed to significantly reduce the memory required for training Mixture-of-Experts (MoE) models. This optimizer achieves a 97.4% reduction in optimizer state memory by employi…