A user has implemented the Qwen 3.5 large language model on field-programmable gate arrays (FPGAs) using INT4 quantization. The project involved utilizing relatively inexpensive eBay mining hardware, specifically SQRL FK33 and Jungle Cat boards, to achieve LLM inference. Initial tests with the Qwen 3.5 9B model demonstrated inference speeds of approximately 2 tokens/second at 75MHz, with projections for the 27B model indicating significantly higher performance with multi-card setups and optimized RTL. AI
IMPACT Demonstrates novel hardware implementations for LLMs, potentially enabling more accessible and cost-effective inference solutions.
RANK_REASON User-implemented LLM on FPGA hardware, not a frontier release from a major lab.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →