Perplexity has open-sourced Lily, a local inference engine designed for hybrid compute within its Perplexity Computer product. Lily is specifically optimized for running Qwen3.6-35B-A3B models on Apple silicon, treating it as a distinct platform. Benchmarks show Lily outperforms MLX-LM in both prefill and decode throughput on an M5 Max MacBook Pro, maintaining output quality while efficiently handling the different computational demands of prefill and decode workloads. AI
IMPACT Enables more efficient local LLM inference on Apple Silicon, potentially improving hybrid compute performance for applications like Perplexity Computer.
RANK_REASON Open-sourcing of a specialized inference engine for local hardware.
- Apple Inc.
- Apple Silicon
- Lily
- M5 Max MacBook Pro
- mlx-lm
- Perplexity
- Perplexity Computer
- Qwen3.6-35B-A3B
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →