A user encountered significant issues with Ollama's prompt handling, where the `num_ctx` setting silently truncated prompts longer than 2048 tokens. This led to a drastic decrease in accuracy for their local RAG bot, dropping from 81% to 46%. The truncation caused the model to miss crucial instructions and top-ranked retrieved chunks, resulting in incorrect JSON formatting and factually wrong answers. The problem was resolved by explicitly setting `num_ctx` to a higher value (e.g., 8192) in the Modelfile or via the native API, which restored accuracy to 81% at the cost of increased VRAM usage. AI
IMPACT Highlights potential pitfalls in local LLM deployment and prompt engineering, impacting developers using Ollama for RAG applications.
RANK_REASON User-reported issue with a specific software tool's functionality.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →