half-precision floating-point format
PulseAugur coverage of half-precision floating-point format — every cluster mentioning half-precision floating-point format across labs, papers, and developer communities, ranked by signal.
- instance of Int8 70%
- used by Int8 70%
- used by graphics processing unit 70%
- instance of Int4 70%
- instance of bfloat16 70%
- instance of single-precision floating-point format 70%
- instance of Fp8 70%
- instance of INTS8 70%
- used by Flashattention 70%
- used by central processing unit 60%
- competes with INTS8 60%
- competes with Int8 50%
5 day(s) with sentiment data
-
New checksum method enhances CNN fault detection on edge devices
Researchers have developed a new lightweight fault-detection technique called Carry-Through Checksum for convolutional neural networks (CNNs) used in edge applications. This method embeds filters into convolutional laye…
-
Qwen3-8B model scaled for ultra-low-bit language processing
Researchers have successfully scaled post-training ternarisation techniques to the Qwen3-8B language model, aiming to reduce storage and memory requirements. The study involved a comprehensive evaluation, including repr…
-
Embedding table precision is key for LLM size reduction
A recent experiment explored the impact of quantization on Transformer models, revealing that the embedding table constitutes a significant portion (72%) of the model's parameters. The research found that quantizing the…
-
New method enhances 2-bit LLM weight decoding efficiency
Researchers have developed a novel multi-shell decoding method for 2-bit LLM weights, aiming to improve efficiency and quality. The proposed approach includes an offline expansion into GPU layouts and a fused dequantize…
-
LinkedIn unveils new GPU retrieval for semantic search
Researchers have developed a new retrieval framework for semantic search on LinkedIn, aiming to improve the relevance of profile suggestions for users. The system partitions embeddings into eight category-supervised seg…
-
New research reveals quantization-triggered backdoors in LLMs
A new research paper details a security vulnerability in large language models (LLMs) where backdoors can be triggered by post-training quantization. The study formalizes this issue through Quantization Behavioral Equiv…
-
New framework enhances attention retrieval in hyperbolic spaces
Researchers have developed HCC+, a theoretical framework for improving attention retrieval in hyperbolic spaces. This new method addresses the lack of deterministic guarantees in existing techniques regarding attention-…
-
AQLoRA offers faster quantized fine-tuning for large language models
Researchers have developed AQLoRA, a novel method for faster quantized fine-tuning of large language models. This technique optimizes the balance between memory savings and training speed, which is a known limitation of…
-
Guide Explains VRAM Needs for Local LLM Deployment
Running large language models locally requires careful VRAM management, as model size and quantization significantly impact memory usage. While there's no exact formula, VRAM needs can be estimated by considering the mo…
-
Developer creates ultra-lightweight LLM runnable on CPUs
A developer has created a custom quantized large language model, SHADOW-250M-Instruct, with 250 million parameters trained on 30 billion tokens. This model is designed for extreme efficiency, deploying at just 60 MB and…
-
Llama 3.2:1b quantization levels show predictable memory scaling but similar response quality
A comparison of quantization levels for the Llama 3.2:1b model revealed that memory usage scales predictably with bit-width, with Q4, Q8, and FP16 variants consuming approximately 0.94 GB, 1.24 GB, and 2.57 GB respectiv…
-
Chenghengwei integrates NPU and GPGPU in new edge AI chip
Chenghengwei has launched its first flagship SoC, the CH3715, which integrates both an NPU and a GPGPU. This decision deviates from the industry trend of prioritizing NPUs for edge AI chips. The company's CEO, Chu Libin…
-
Peking University and StepFun unveil TensorCast for faster LLM inference · 2 sources tracked
Researchers from Peking University and StepFun have developed TensorCast, a new programmable tensor management layer designed to optimize large language model infrastructure. This system aims to significantly reduce the…
-
AI models shrink via quantization and pruning for efficiency
Quantization and pruning are techniques used to reduce the size and computational requirements of large AI models like ChatGPT and Midjourney. These methods decrease the precision of the numbers representing model weigh…
-
Hand-written PTX kernels show significant speedups for INT8/INT4 GEMM on NVIDIA L4 GPUs
A new research paper explores the performance benefits of using hand-written PTX (Parallel Thread Execution) kernels for GEMM (General Matrix Multiply) operations on NVIDIA L4 GPUs, compared to the standard WMMA (Warp M…
-
Quantized LLMs face 'alignment collapse,' erasing safety guardrails
Post-training quantization (PTQ) of large language models (LLMs) can lead to a phenomenon called "alignment collapse," where safety guardrails like RLHF and DPO are silently erased when models are compressed to lower bi…
-
AI models may ditch matrix multiplication for addition-only hardware
Researchers are exploring a shift from traditional matrix multiplications in AI models to simpler addition-only operations, aiming to overcome the memory bandwidth bottleneck. This approach, which involves using extreme…
-
LLM Audits Miss 90% of Safety Failures Due to Post-Training Compression
A new analysis suggests that standard auditing practices for large language models (LLMs) fail to detect significant safety failures that emerge after model compression. Compressing full-precision models to lower bit-wi…
-
AirLLM enables 70B models on 4GB GPU via layer-wise inference · 8 sources tracked
The open-source project AirLLM has gained significant traction, reaching over 27,000 stars on GitHub. Its core innovation allows large language models, specifically 70 billion parameter models, to run on a single 4GB GP…
-
INT4 Weight-Only Quantization: Decode Speedup, Prefill Stagnation Explained
Weight-only INT4 quantization, while effective for reducing memory traffic and speeding up the decoding phase of LLM inference, does not improve the prefill phase. This is because prefill is compute-bound, meaning it is…