Researchers have developed a new method for lossless source code compression using Large Language Models (LLMs) that improves upon existing techniques. This approach, called Thresholded Symbol Ranking, bounds LLM predictions to the top-ranked symbols, storing exceptions separately. Evaluations across 30 LLMs show this method achieves up to 37% better compression ratios and is 40% faster than previous LLM-based compressors. While still slower than general-purpose compressors like zstd, it offers a significantly improved compression gain of up to 82% relative to them, demonstrating LLMs' effectiveness in capturing source code's unique regularities. AI
IMPACT Offers a new trade-off point between compression ratio and speed for large-scale code archives, potentially impacting storage costs and data transfer efficiency.
RANK_REASON Academic paper detailing a novel method for source code compression using LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- bzip2
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
- Scite
- Software Heritage
- Zstandard
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →