Perplexity has open-sourced Lily, a specialized inference engine built with Rust and Metal for running the Qwen3.6-35B-A3B model on Apple Silicon. This engine is designed for high performance by tightly integrating model structure, execution plans, and kernel selection, bypassing traditional frameworks like PyTorch and MLX. Lily achieves significant speedups in both prefill and decode operations compared to existing implementations, particularly for long contexts on Macs with substantial unified memory. AI
IMPACT Enables more efficient local LLM inference on Apple hardware, potentially improving user experience for Perplexity's products.
RANK_REASON Open-sourcing of a specialized inference engine for a specific model and hardware.
- Apple Silicon
- Lily
- macOS
- Metal
- Mlx
- mlx-lm
- OpenAI
- Perplexity
- Perplexity Computer
- PyTorch
- Qwen3.6-35B-A3B
- Rust
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →