Perplexity has open-sourced Lily, a specialized inference engine built with Rust and Metal for running the Qwen3.6-35B-A3B model on Apple Silicon. This engine is designed for narrow hardware optimization, achieving up to 1.35x faster decode and 1.23x faster prefill speeds compared to general-purpose frameworks like MLX-LM. Lily's architecture bypasses traditional frameworks like PyTorch and MLX, utilizing hand-written Metal kernels for execution and offering an OpenAI-compatible API. AI
IMPACT Specialized inference engines like Lily could accelerate local AI model deployment on consumer hardware by optimizing performance beyond general-purpose frameworks.
RANK_REASON Perplexity open-sourced a specialized inference engine, Lily, which is a tool for running a specific model on specific hardware.
Read on Mastodon — mastodon.social →
- Apple Silicon
- Lily
- macOS
- Metal
- Mlx
- MLX-LM
- OpenAI
- Perplexity
- Perplexity Computer
- PyTorch
- Qwen3.6-35B-A3B
- Rust
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →