Apple M1 Max
PulseAugur coverage of Apple M1 Max — every cluster mentioning Apple M1 Max across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
Ollama production update deferred due to past issues and shared resource risks
The author decided against updating Ollama to version 0.32.0 on their production machine due to potential risks, despite the release notes advertising a new "agent" feature. The decision was informed by past incidents w…
-
Ollama 0.30.8 on Apple Silicon: MLX runner not active for GGUF models
A recent analysis of Ollama version 0.30.8 revealed that despite the binary containing code for an MLX runner, it does not appear to be utilized when running standard GGUF models on Apple Silicon. The investigation, con…
-
Qwen3.5-9B model's "thinking" tokens inflate output, slowing local LLM performance
A user tested the Qwen3.5-9B model on an Apple M1 Max with 64GB of RAM, using Ollama for local execution. While the model's prompt suggested it could outperform GPT-4 in Japanese, the test focused on actual performance …
-
Developer builds agent-orchestra to run ChatGPT, Claude, and local LLMs in parallel
A developer has created a system called agent-orchestra to manage multiple large language models (LLMs) concurrently for coding tasks. This tool allows parallel execution of tasks across models like ChatGPT, Claude, and…
-
Qwen3 LLM runs up to 4.52x faster on Apple Silicon with ExecuTorch MLX delegate
A recent technical exploration demonstrates significant speed improvements when running the Qwen3-0.6B language model on Apple Silicon using ExecuTorch's experimental MLX delegate. The MLX delegate, which leverages Appl…
-
M1 Max LLM Benchmark: Larger MoE Models Prove Faster Locally
A local LLM benchmark on an Apple M1 Max with 64GB of RAM revealed that larger models are not always slower. The test, using Ollama, found that a 23.9GB Qwen3.6 MoE model achieved 60.4 tokens/sec, outperforming a smalle…
-
GLM-5.2 slower than Qwen3.6 on 64GB Mac due to RAM limits and active parameters · 1 source tracked
A comparison of the GLM-5.2 and Qwen3.6 large language models on a 64GB Mac revealed that GLM-5.2 is significantly slower, contrary to some claims. The primary reasons are that GLM-5.2, an open-source 753B parameter Mix…
-
M1 Max inference engines benchmarked: rapid-mlx leads
A hobbyist benchmarked several inference engines on an M1 Max MacBook Pro using the Qwen3.5-4B model. The results, submitted to the mlx-chronos community benchmark, indicate that rapid-mlx offers the best performance in…
-
Phosphene AI video tool adds LoRA support, runs on Macs with 16GB RAM
The open-source AI video generation tool Phosphene has rapidly updated with LoRA support and CivitAI integration, allowing users to apply custom LoRA models like Retro anime LoRA. Additionally, tips have emerged for run…
-
AI model predicts stuttering events from audio, deploys on-device
Researchers have developed a new Convolutional Neural Network (CNN) model capable of predicting upcoming stuttering events from short audio clips. The 616K-parameter model, trained on the SEP-28k dataset, demonstrates a…
-
Consumer-grade graphics cards can quickly get started! MiniCPM-o 4.5 from Mianbi Intelligent releases technical report
MiniCPM-o 4.5 is a new 9B parameter omni-modal large language model designed for real-time, full-duplex interaction. It can simultaneously process and generate audio, video, and text, enabling proactive behaviors and co…