A user on r/LocalLLaMA is experiencing a crash when attempting to run the GLM-5.2 model with contexts larger than 8k, despite the model being trained for up to 1 million tokens. The crash occurs with a fatal error in llama-sampling.cpp, resulting in all logit probabilities being NaN. The user has tried various troubleshooting steps, including adjusting DSA settings, KV cache formats, and pulling the latest commits, but the issue persists, suggesting a potential bug in the DSA or indexer path for long contexts. AI
IMPACT Users are encountering limitations with long context windows in the GLM-5.2 model, hindering its full potential for advanced applications.
RANK_REASON User is reporting a bug with a specific model and tool, seeking a workaround.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →