实体
VKUE
VKUE
PulseAugur coverage of VKUE — every cluster mentioning VKUE across labs, papers, and developer communities, ranked by signal.
总计 · 30天
0
90 天内 2
发布 · 30天
0
90 天内 0
论文 · 30天
0
90 天内 0
层级分布 · 90 天
主题
最近 · 第 1/1 页 · 共 2 条
-
VIDRAFT 推出双 LLM 服务引擎,兼顾 GPU 吞吐量和 CPU 覆盖范围
VIDRAFT 开发了两款不同的语言大模型服务引擎,分别针对不同的优化目标。VKAE 是一款内核级加速引擎,旨在最大化 GPU 上的吞吐量,在多请求场景下性能提升高达 23.4 倍,速度超过 10,000 tokens/秒。相比之下,VKUE 通过优化内存带宽而非原始计算能力,使得一个 34.7B 参数的模型能够在包括 CPU 在内的各种硬件上运行,适用于受监管或本地部署的工作负载。
-
Tencent and VIDRAFT showcase sparse MoE models with reduced active parameters
Tencent has released Hy3, a 295-billion-parameter Mixture-of-Experts (MoE) model that utilizes only 21 billion active parameters per forward pass, significantly reducing inference costs. This MoE architecture, featuring…