PulseAugur
EN
LIVE 08:28:45

Ling-3.0-tiny model runs on 8GB NVIDIA Orin Nano with 128K context

A user has successfully run the Ling-3.0-tiny model on an NVIDIA Orin Nano Super 8GB device, achieving a 128K context window with IQ4_NL quantization. The model, which utilizes a hybrid KDA+MLA reasoning MoE architecture, fits within the 8GB of unified RAM and demonstrates promising performance for its hardware class. Despite some degradation in retrieval accuracy at the maximum context length, the setup is considered viable for agent hosting and handling simpler tasks. AI

IMPACT Demonstrates the feasibility of running advanced LLMs with large context windows on low-power, edge devices.

RANK_REASON User successfully runs a specific model on consumer hardware, demonstrating its capabilities and limitations.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Ling-3.0-tiny model runs on 8GB NVIDIA Orin Nano with 128K context

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Puzzleheaded_Base302 ·

    Ling-3.0-tiny is a very interesting model. Run on NVIDIA Orin Nano Super 8GB at 128K context with IQ4_NL quant.

    <!-- SC_OFF --><div class="md"><p>I have been searching for suitable model to run on my 8GB RAM toy, NVIDIA Orin Nano Super 8GB. This little toy was priced at $249 earlier this year (not any more), and pulls very little power when idle. It was an interesting device that suitable …