PulseAugur
EN
LIVE 23:39:57

Perplexity open-sources Lily inference engine for Apple Silicon

Perplexity has open-sourced Lily, a specialized inference engine built with Rust and Metal for running the Qwen3.6-35B-A3B model on Apple Silicon. This engine is designed for narrow hardware optimization, achieving up to 1.35x faster decode and 1.23x faster prefill speeds compared to general-purpose frameworks like MLX-LM. Lily's architecture bypasses traditional frameworks like PyTorch and MLX, utilizing hand-written Metal kernels for execution and offering an OpenAI-compatible API. AI

IMPACT Specialized inference engines like Lily could accelerate local AI model deployment on consumer hardware by optimizing performance beyond general-purpose frameworks.

RANK_REASON Perplexity open-sourced a specialized inference engine, Lily, which is a tool for running a specific model on specific hardware.

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

Perplexity open-sources Lily inference engine for Apple Silicon

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Perplexity open-sourced a specialized inference engine, Lily, which is a tool for running a specific model on specific hardware.
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
12 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [3]

  1. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon

    <p>Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one model on one chip family, it averages 1.23x MLX-LM's prefill throughput and 1.35x its decode throughput on a 40-core, 128 …

  2. r/LocalLLaMA TIER_1 English(EN) · /u/saltexx ·

    We open-sourced Paddock, our Rust/C++ inference engine with its own CUDA kernels (MIT/Apache-2.0)

    <!-- SC_OFF --><div class="md"><p>I'm one of the developers. We said in August it would go open source in September and it did last night. MIT or Apache-2.0, pick one. The repo you see is our internal repo, kernels included, so from now on everything happens in public.</p> <p>It'…

  3. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Perplexity has open sourced Lily, a Rust and Metal inference engine for running Qwen3.6-35B-A3B on Apple Silicon. The specialised engine achieves 1.23x faster p

    Perplexity has open sourced Lily, a Rust and Metal inference engine for running Qwen3.6-35B-A3B on Apple Silicon. The specialised engine achieves 1.23x faster prefill and 1.35x faster decode than MLX-LM, demonstrating how narrow hardware optimisation can beat general-purpose fram…