Ollama version 0.40 introduces new metrics, including a prefill rate and decision model probabilities, though the prefill rate calculation is flawed. The prefill rate incorrectly inflates the speed by factoring in cached tokens, leading to misleadingly high numbers. Decision models, such as nimble and tev1, provide probabilities for choices or scores instead of generating text, but can be mistakenly offered for chat completions if not properly filtered by capability. AI
IMPACT This update to Ollama provides developers with more detailed performance metrics, though users should be aware of potential inaccuracies in the reported prefill rate.
RANK_REASON The item details a new version of a local LLM runner and its new features, including a flawed metric calculation.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →