PulseAugur
EN
LIVE 23:23:11

New SAOD Technique Compresses 744B LLM from 1.5TB to Under 100GB

A new technique called Session-Adaptive Orthogonal Distillation (SAOD) has been proposed to significantly compress large language models. This method aims to reduce the size of a 744 billion parameter model, which currently occupies 1.5 terabytes, to under 100 gigabytes. The potential impact is that models of this scale could potentially run on hardware with as little as 8GB of VRAM. AI

IMPACT Enables running significantly larger models on consumer-grade hardware, democratizing access to advanced AI capabilities.

RANK_REASON The item describes a new technical method for model compression, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New SAOD Technique Compresses 744B LLM from 1.5TB to Under 100GB

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    Session-Adaptive Orthogonal Distillation (SAOD)? Technology compresses 744B (1.5TB) to under 100GB?

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v3shir/sessionadaptive_orthogonal_distillation_saod/"> <img alt="Session-Adaptive Orthogonal Distillation (SAOD)? Technology compresses 744B (1.5TB) to under 100GB?" src="https://preview.redd.it/tkgryjv1cueh1…