Sram
PulseAugur coverage of Sram — every cluster mentioning Sram across labs, papers, and developer communities, ranked by signal.
9 day(s) with sentiment data
-
AMD acquires Taalas to push specialized AI inference chips
AMD has acquired Taalas, a Canadian AI inference chip company, to integrate its specialized hardware into AMD's accelerator roadmap. Taalas focuses on creating chips optimized for specific AI models, aiming to reduce in…
-
AI models etched into silicon for faster inference, bypassing memory bottlenecks
A new trend is emerging where AI models are being etched directly into silicon, a process that permanently embeds model weights as physical transistors. This approach, exemplified by companies like Taalas and previously…
-
SK Hynix, SanDisk unveil High Bandwidth Flash for AI inference memory wall
SK Hynix and SanDisk have collaborated to develop High Bandwidth Flash (HBF), a new memory tier designed to address the memory wall challenges in AI inference. HBF places large-capacity NAND flash memory close to variou…
-
MCP, eMMC, and eMCP: Understanding Integrated Memory Technologies
This article delves into the distinctions between Multi-Chip Package (MCP), embedded MultiMediaCard (eMMC), and embedded Multi-Chip Package (eMCP) technologies. MCP is a packaging method that integrates multiple chips l…
-
DJI-spinoff Amflow e-bikes surpass 1B yuan revenue, challenging market leaders
Amflow, an e-bike company spun out of DJI, has achieved over 1 billion yuan in revenue, selling high-end electric mountain bikes priced significantly higher than typical market offerings. The company, along with its sis…
-
New optical receiver updates robot AI via light, cutting energy use
Researchers have developed a novel optical receiver that can directly alter a robot's AI model parameters using light. This technology aims to overcome the significant memory and energy demands of current AI systems by …
-
NumPy implementation of Flash Attention demonstrates significant memory savings
This article details the implementation of Flash Attention from first principles using NumPy. Flash Attention optimizes transformer models by avoiding the materialization of large N×N attention score matrices, which con…
-
New UEP codec slashes AI inference memory costs by up to 62.5%
Researchers have developed a new method for protecting memory in AI inference by analyzing bit-position fault sensitivity in various models and floating-point formats. They found that certain lower-order bits have minim…
-
New ThRIve method boosts CNN inference robustness in PIM architectures
Researchers have developed ThRIve, a novel training methodology designed to enhance the thermal robustness of Convolutional Neural Network (CNN) inference on Processing-In-Memory (PIM) architectures. This approach utili…
-
3D Gaussian renderer implemented on AI accelerator with on-chip SRAM
Researchers have developed the first implementation of a 3D Gaussian renderer specifically for an Intelligence Processing Unit (IPU), a type of AI accelerator. This system utilizes 1,472 independent tiles with on-chip S…
-
36Kr: Jiangling Motors sales up 0.04% in June; Beijing Junzheng sees DRAM price hikes
36Kr reported that Jiangling Motors sold 35,742 vehicles in June, a slight year-over-year increase of 0.04%. Year-to-date sales reached 189,000 units, up 9.46%. Separately, Beijing Junzheng noted significant price incre…
-
Beijing Junzheng forecasts continued DRAM price hikes amid tight supply
Beijing Junzheng anticipates further price increases for DRAM in the third quarter, with a possibility of adjustments in the fourth quarter due to continued tight supply. The company also noted a price increase for some…
-
New framework enhances fault tolerance in FPGA-based CNN accelerators
Researchers have developed ProWAFT, a novel fault-tolerance framework designed for CNN accelerators implemented on SRAM-based FPGAs. This system addresses the challenge of transient faults that can compromise reliabilit…
-
Cluster-Scale Memory Introduced to Tackle AI Chip Bottlenecks
Cluster-Scale Memory (CSM) has been introduced to address low-latency workload challenges in AI chips. Current AI chips utilizing High Bandwidth Memory (HBM) face limitations in achieving SRAM-level decode speeds becaus…
-
Li Auto aims for Tesla FSD V14 parity with self-developed AI chips and models
Li Auto is developing its autonomous driving capabilities to match Tesla's FSD V14, focusing on safety, efficiency, and comfort, alongside advanced features like recognizing special vehicles and traffic police signals. …
-
Flash Attention Mechanics Explained: Tiled Attention in SRAM
This article delves into the mechanics of Flash Attention, a technique designed to optimize the self-attention mechanism in AI models. It explains how tiled attention, a method for processing attention computations in s…
-
SRAM Supply Contradiction Questioned Amidst Logic Wafer Constraints
A conversation on X highlights a perceived contradiction regarding SRAM supply. The dialogue questions the availability of SRAM, given that the logic wafers used in its fabrication are reportedly supply-constrained. Thi…
-
Qualcomm unveils near-memory AI architecture to boost performance
Qualcomm has introduced a new near-memory AI architecture called High Bandwidth Compute (HBC) designed to overcome the memory wall limitations in AI workloads. This architecture places AI accelerators directly beneath L…
-
Groq LPU gains traction in AI inference, challenging GPU dominance
Groq's Language Processing Unit (LPU) is gaining traction in the AI inference market, moving beyond niche applications to become a recognized component in AI infrastructure. This shift is driven by the increasing demand…
-
SNIA launches MRAM SIG to standardize interfaces and boost adoption
The Storage Networking Industry Association (SNIA) has launched a Magnetoresistive Random-Access Memory (MRAM) Special Interest Group (SIG) to foster MRAM adoption. This group aims to standardize MRAM technologies and d…