Tilert
PulseAugur coverage of Tilert — every cluster mentioning Tilert across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
NVIDIA GPUs achieve 500 tokens/sec inference with TileRT engine
NVIDIA GPUs, when utilized with TileRT's persistent engine, can achieve inference speeds of 500 tokens per second per user. This configuration reportedly surpasses the performance of traditional setups on identical hard…
-
NVIDIA's TileRT InferenceX challenged by specialized AI hardware
SemiAnalysis is questioning whether NVIDIA's TileRT InferenceX software, running on its GPUs, can compete with specialized AI hardware from companies like Cerebras, Groq, and SambaNova. The software is designed for ultr…
-
Xiaomi claims 1000+ tokens/sec on 1T parameter model with 8 GPUs
Xiaomi's MiMo team has announced MiMo-V2.5-Pro UltraSpeed, a 1 trillion parameter Mixture-of-Experts model capable of exceeding 1,000 tokens per second. This performance was achieved on a standard 8-GPU server, utilizin…
-
Xiaomi achieves 1000 tokens/sec on 1T-parameter model with commodity GPUs
Xiaomi's MiMo team has released MiMo-V2.5-Pro-UltraSpeed, a new inference mode for their 1-trillion-parameter model that achieves over 1000 tokens per second on commodity GPUs. This significant speedup is attributed to …
-
Zhipu AI launches GLM-5.1-highspeed API at 400 tokens/s
Zhipu AI has released GLM-5.1-highspeed, a new API for its GLM-5.1 model that achieves an inference speed of 400 tokens per second. This new offering is positioned as the fastest among leading global LLM providers and h…