PulseAugur
实时 21:12:24
实体 anemll-flash-llama.cpp

anemll-flash-llama.cpp

PulseAugur coverage of anemll-flash-llama.cpp — every cluster mentioning anemll-flash-llama.cpp across labs, papers, and developer communities, ranked by signal.

Show in brief
总计 · 30天
1
90 天内 1
发布 · 30天
0
90 天内 0
论文 · 30天
0
90 天内 0
层级分布 · 90 天
主题
情绪 · 30 天

1 天有情绪数据

最近 · 第 1/1 页 · 共 1 条
  1. TOOL · CL_190629 ·

    Flash-MoE 技术让大型 AI 模型能在 16GB 内存的 Mac 上运行

    一种名为 anemll-flash-llama.cpp 的新技术,能够让大型专家混合(MoE)模型在内存仅为 16GB 的 Mac 上运行。该方法将模型专家存储在 SSD 上,仅将必要的专家加载到小型缓存中,从而显著降低内存需求。基准测试表明,虽然这种方法受存储 I/O 限制,但它使得 Qwen3.5-35B-A3B 等模型能在消费级硬件上使用,其中 Q3_K_M 等量化方法被证明更实用。