PulseAugur
EN
LIVE 07:46:48
ENTITY AMD MI308X

AMD MI308X

PulseAugur coverage of AMD MI308X — every cluster mentioning AMD MI308X across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
0
6 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
0 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
RECENT · PAGE 1/1 · 6 TOTAL
  1. TOOL · CL_212517 ·

    Mingxin FX100 storage boosts AI inference, aiding domestic substitution

    Mingxin's FX100 storage solution offers significant performance improvements for AI inference, particularly in domestic substitution efforts within the Xinchuang environment. By focusing on the storage protocol and data…

  2. TOOL · CL_206773 ·

    LLM TTFT Reduction Strategies Explored Across Compute, Memory, and Storage

    Reducing the first-token latency (TTFT) of large language models is crucial for user experience and performance. This involves optimizing four key areas: compute, GPU memory, storage, and overall architecture. Technique…

  3. TOOL · CL_181549 ·

    Mingxin FX100 boosts LLM inference with KV Cache reuse · 2 sources tracked

    Mingxin FX100 has demonstrated significant performance improvements in multi-turn dialogue scenarios for large language models. By implementing KV Cache reuse strategies, which involve caching key-value tensors from pre…

  4. TOOL · CL_179717 ·

    AI inference cards slash database query latency by up to 32%

    A new study highlights how domestic AI inference acceleration cards, specifically the Mingxin FX100, can significantly improve real-time database query performance. By optimizing storage access paths and reducing model …

  5. TOOL · CL_174738 ·

    KV Cache tiering boosts LLM inference speed and cuts costs

    A new approach to managing KV Cache in large language model inference suggests treating it as a high-frequency access subset within the warm storage tier, rather than in the traditional hot or cold tiers. This strategy,…

  6. TOOL · CL_174551 ·

    Mingxin Tech boosts GPU utilization by 16% via optimized model switching

    Mingxin Technology has demonstrated significant improvements in GPU compute utilization by addressing model switching and cold-start latency. Through a three-step optimization process involving tiered KV Cache accelerat…