PulseAugur
EN
LIVE 07:33:30

Perplexity open-sources Lily inference engine for Apple Silicon

Perplexity has open-sourced Lily, a specialized inference engine built with Rust and Metal for running the Qwen3.6-35B-A3B model on Apple Silicon. This engine is designed for high performance by tightly integrating model structure, execution plans, and kernel selection, bypassing traditional frameworks like PyTorch and MLX. Lily achieves significant speedups in both prefill and decode operations compared to existing implementations, particularly for long contexts on Macs with substantial unified memory. AI

IMPACT Enables more efficient local LLM inference on Apple hardware, potentially improving user experience for Perplexity's products.

RANK_REASON Open-sourcing of a specialized inference engine for a specific model and hardware.

Read on MarkTechPost →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Perplexity open-sources Lily inference engine for Apple Silicon

How we ranked this

Signal score
42 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Open-sourcing of a specialized inference engine for a specific model and hardware.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon

    <p>Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one model on one chip family, it averages 1.23x MLX-LM's prefill throughput and 1.35x its decode throughput on a 40-core, 128 …