Ollama has released version 0.34.1, introducing several key updates. The release makes MLX safetensors "ollama create" functionality no longer experimental and improves memory handling for MLX on Apple Silicon. Additionally, GGUF model creation now necessitates the use of llama.cpp tooling for safetensor conversion and quantization, and runaway repeat token detection has been refined to reduce false positives. AI
IMPACT Improves the usability and performance of local LLM deployment for users of Ollama.
RANK_REASON This is a software release for a tool that facilitates running LLMs locally, not a frontier model release or core research.
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →