PulseAugur
EN
LIVE 23:53:21

Meta's custom AMD chip sacrifices LLM performance for recsys optimization

SemiAnalysis reports that Meta is developing a custom AMD MI400-series chip, which is half the size of the standard MI455X and optimized for recommendation system workloads and memory bandwidth efficiency. This custom chip uses significantly less High Bandwidth Memory (HBM) than the standard version, making it less suitable for LLM inference and training. The report also criticizes Meta's broader infrastructure strategy, citing past issues with the GB200 NVL72 Ariel and suggesting that the custom chip design may hinder Meta's ability to rent out compute resources, a strategy inspired by Elon Musk's approach. AI

IMPACT Meta's custom chip design may limit its LLM capabilities, potentially impacting its competitiveness in AI development.

RANK_REASON The cluster consists of multiple tweets from SemiAnalysis discussing Meta's custom chip design and infrastructure strategy, offering analysis and criticism rather than a primary announcement.

Read on X — SemiAnalysis →

AI-generated summary · Google Gemini · from 7 sources. How we write summaries →

Meta's custom AMD chip sacrifices LLM performance for recsys optimization

COVERAGE [7]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @elonmusk We think that the MI455X will be a great chip as long as @AnushElangovan invests enough in software and automated testing capabilities to fix AMD's lo

    @elonmusk We think that the MI455X will be a great chip as long as @AnushElangovan invests enough in software and automated testing capabilities to fix AMD's long history of poor software quality, but Meta's overengineering leads to less flexibility in the overall strategy that A…

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Another issue that Meta's custom half-size package design will cause is that it will be harder for Zuck to rent them out, following his strategy of copying @elo

    Another issue that Meta's custom half-size package design will cause is that it will be harder for Zuck to rent them out, following his strategy of copying @elonmusk's neocloud strategy. 6/7🧵

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Another example of Meta's overengineering obsession is GB200 NVL72 Ariel, which caused massive infrastructure issues with its cross-rack NVLink ACC cables due t

    Another example of Meta's overengineering obsession is GB200 NVL72 Ariel, which caused massive infrastructure issues with its cross-rack NVLink ACC cables due to signal integrity problems, as Meta was obsessed with having the Grace CPU at a 1:1 ratio with the GPU. This decision h…

  4. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    This is yet another instance of the recsys infrastructure strategy team at Meta wanting to look like they “add value” by doing weird micro-optimizations that hu

    This is yet another instance of the recsys infrastructure strategy team at Meta wanting to look like they “add value” by doing weird micro-optimizations that hurt other important divisions at Meta. 4/7🧵

  5. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    But the issue is that Meta’s custom MI400-series SKU is not as optimized for LLM inference and training. The decision was made before TBD Lab was formed or coul

    But the issue is that Meta’s custom MI400-series SKU is not as optimized for LLM inference and training. The decision was made before TBD Lab was formed or could have its say. Given the significant decreases in compute and HBM in its custom SKU, it will be less attractive to

  6. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Compared to a normal MI455X package, it will use six HBM4 8i stacks instead of 12 HBM4 12Hi stacks. The reasoning is that Meta’s recsys infrastructure strategy

    Compared to a normal MI455X package, it will use six HBM4 8i stacks instead of 12 HBM4 12Hi stacks. The reasoning is that Meta’s recsys infrastructure strategy wanted to have a CPU compute-to-GPU compute ratio tuned for recsys and to optimize memory $/BW. 2/7🧵 https://t.co/jtfrdP…

  7. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    ALERT🚨🚨: META's CUSTOM AMD MI400-series chip will be half the size of a normal MI455X chip. It is "optimized" for recsys workloads and $/Memory Bandwidth. It wi

    ALERT🚨🚨: META's CUSTOM AMD MI400-series chip will be half the size of a normal MI455X chip. It is "optimized" for recsys workloads and $/Memory Bandwidth. It will use ~144GB of HBM instead of 432GB. We break it down below👇️ 1/7🧵 https://t.co/hiz53VYqkv