A user on Reddit's r/LocalLLaMA subreddit reported achieving 600 tokens per second for single requests using the qwen3.6 35ba3b model. This performance was achieved with the NInfer framework running on an Nvidia RTX Pro 6000 workstation. The user described the model as a capable, albeit not top-tier, option for tasks like information retrieval and code generation, noting its speed makes it viable even for more token-intensive operations. AI
IMPACT Demonstrates significant speed improvements for local LLM inference on high-end hardware.
RANK_REASON User report on hardware/software performance for a specific model.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →