PulseAugur
EN
LIVE 22:52:27
ENTITY mlx-lm

mlx-lm

PulseAugur coverage of mlx-lm — every cluster mentioning mlx-lm across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
4
7 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
0 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

4 day(s) with sentiment data

RECENT · PAGE 1/1 · 8 TOTAL
  1. TOOL · CL_256244 ·

    LLMeter CLI measures LLM performance on local hardware

    LLMeter is a new command-line interface tool designed to measure the performance of large language models (LLMs) on a user's specific hardware and configuration. Unlike traditional leaderboards that test models on optim…

  2. SIGNIFICANT · CL_247329 ·

    Edge0 releases 35B MoE LLM for low-memory devices

    Edge0 has released a preview of its Edge0-35B-A3B model, a 35 billion parameter sparse Mixture-of-Experts (MoE) large language model designed to run efficiently on devices with limited memory. The model requires under 3…

  3. TOOL · CL_233724 ·

    Perplexity open-sources Lily inference engine for Apple Silicon

    Perplexity has open-sourced Lily, a specialized inference engine built with Rust and Metal for running the Qwen3.6-35B-A3B model on Apple Silicon. This engine is designed for narrow hardware optimization, achieving up t…

  4. TOOL · CL_232829 ·

    Perplexity open-sources Lily inference engine for Apple Silicon

    Perplexity has open-sourced Lily, a local inference engine designed for hybrid compute within its Perplexity Computer product. Lily is specifically optimized for running Qwen3.6-35B-A3B models on Apple silicon, treating…

  5. COMMENTARY · CL_202492 ·

    Apple Silicon LLM inference hampered by fragmented software stack

    Running large language models efficiently on Apple Silicon is hampered by a fragmented and immature software ecosystem. Unlike the integrated optimizations available on NVIDIA's CUDA platform, Apple's platform lacks a u…

  6. TOOL · CL_188516 ·

    Gemma models on Apple Silicon suffer silent cache bug, slowing local AI

    A cache bug affecting Gemma models on Apple Silicon has been identified, causing significant slowdowns in local AI agent performance. The issue stems from Gemma's sliding-window attention mechanism, which, when exceedin…

  7. TOOL · CL_167685 ·

    FusionML boosts Apple Silicon AI inference with CPU-GPU co-execution

    Researchers have developed FusionML, a system that optimizes transformer inference on Apple Silicon by co-executing CPU and GPU operations. By addressing limitations in MLX's scheduler that caused serialization, FusionM…

  8. RESEARCH · CL_70649 ·

    Gemma 4 12B local AI model requires configuration tweaks for optimal performance

    Google's Gemma 4 12B model shows promise for local AI setups, but users report that default configurations in tools like LM Studio can hinder its reasoning capabilities. Specific adjustments to Jinja templates and sampl…