PulseAugur
中
实时 12:03:00
实体 LLM in a flash: Efficient large language model inference with limited memory

LLM in a flash: Efficient large language model inference with limited memory

PulseAugur coverage of LLM in a flash: Efficient large language model inference with limited memory — every cluster mentioning LLM in a flash: Efficient large language model inference with limited memory across labs, papers, and developer communities, ranked by signal.

Show in brief
总计 · 30天
1
90 天内 1
发布 · 30天
0
90 天内 0
论文 · 30天
0
90 天内 0
层级分布 · 90 天
主题
关系
情绪 · 30 天

1 天有情绪数据

最近 · 第 1/1 页 · 共 1 条
  1. TOOL · CL_278870 ·

    Flash-MoE 使 397B 参数 LLM 能够在 48GB 笔记本电脑上运行

    一种名为 Flash-MoE 的新技术允许一个拥有 3970 亿参数的庞大模型 Qwen3.5-397B-A17B 在只有 48GB RAM 的 MacBook Pro 等消费级硬件上运行。这是通过利用专家混合(MoE)架构实现的,其中每个 token 只激活模型参数的一小部分,其余参数可以从 SSD 流式传输。这种方法受到一篇先前未发布的 Apple 论文的启发,显著减小了内存占用,尽管与基于云的解决方案相比,推理速度较慢。