knowledge distillation
PulseAugur coverage of knowledge distillation — every cluster mentioning knowledge distillation across labs, papers, and developer communities, ranked by signal.
- used by large-language models 90%
- uses large-language models 90%
- instance of distilled AI model 90%
- developed by large-language models 70%
- developed large-language models 70%
- used by CIFAR-100 70%
- instance of CIFAR-100 70%
- used by alphaXiv 70%
- used by ScienceCast 70%
- used by Gotit.pub 70%
- used by distilled AI model 70%
- used by DagsHub 60%
7 day(s) with sentiment data
-
FLoKD framework enables federated LLM training with reduced communication costs
Researchers have developed FLoKD, a novel framework for federated learning of large language models (LLMs) over wireless networks. This approach utilizes adaptive knowledge distillation by transmitting intermediate LoRA…
-
New backdoor attack targets AI knowledge distillation process
Researchers have demonstrated a new method for backdooring image knowledge distillation, a process typically used to transfer capabilities from large AI models to smaller ones. The attack involves poisoning the distilla…
-
Knowledge Distillation: Shrinking LLMs for Efficient Deployment
Knowledge distillation is a technique used to compress large language models (LLMs) by transferring the learned behaviors of a large "teacher" model into a smaller "student" model. This process is crucial for deploying …
-
New theory explains knowledge distillation in decentralized AI networks
Researchers have developed a convergence theory for knowledge distillation within asynchronous peer-to-peer gossip learning networks. This approach addresses the challenge of averaging models with different parameter co…
-
New research identifies "trait-direction drift" as mechanism for subliminal AI learning
Researchers have identified "subliminal learning" as a mechanism where AI models can unintentionally transfer hidden traits from a teacher model to a student model during knowledge distillation. This phenomenon, termed …
-
Study questions environmental benefits of knowledge distillation in machine translation
A new study published on arXiv investigates the environmental impact of knowledge distillation (KD) in machine translation. Researchers evaluated KD methods using the Machine Learning Life Cycle Assessment tool, conside…
-
New SARA attack bypasses Vision Transformer privacy defenses
A new research paper details a feature inversion attack called SARA that can reconstruct input images from Vision Transformer (ViT) embeddings transmitted in split-inference systems. The attack demonstrates that token s…
-
AI research warns against over-reliance on teacher mimicry metrics
A new research paper analyzes the gap between a student AI model's ability to mimic a teacher model and its actual performance on a task. The study uses a minimal three-party model to demonstrate that while the student'…
-
AI knowledge distillation explained: Smaller models mimic larger ones
Knowledge distillation is a technique in artificial intelligence where a smaller, more efficient model is trained to mimic the behavior of a larger, more complex model. This process allows for the deployment of powerful…
-
Knowledge distillation enables smaller AI models to learn from larger ones
Knowledge distillation (KD) is a technique that allows smaller, more efficient AI models to learn from larger, more capable "teacher" models. Instead of training from scratch on basic labels, a "student" model is traine…
-
New Persistent Cross Entropy Metric Introduced for Topological Data Analysis
Researchers have introduced Persistent Cross Entropy (PCE), a novel method to measure the cross-entropy between two persistence diagrams. This new metric addresses the challenge of differing event spaces in persistence …
-
Knowledge distillation can improve CNNs in data-scarce settings
Researchers have investigated knowledge distillation (KD) for training smaller, more efficient Convolutional Neural Networks (CNNs) by transferring knowledge from larger teacher models. While typically applied at the fi…
-
Diffusion Language Models Advance with New Efficiency and Safety Techniques · 10 sources tracked
Recent research explores advancements in diffusion language models (DLMs), focusing on improving their efficiency, safety, and capabilities. Papers introduce methods like Q-Skew for privacy risk assessment and PII extra…
-
New self-distillation framework enhances lightweight IoT attack detection
Researchers have developed a new teacher-free latent self-distillation framework called Twin Autoencoder (TAE) for lightweight Internet of Things (IoT) attack detection. Unlike traditional knowledge distillation methods…
-
New Adaptive Entropy Distillation method improves LLM knowledge transfer
Researchers have proposed a new knowledge distillation method called Adaptive Entropy Distillation (AED) that aims to improve the transfer of capabilities from large language models (LLMs) to smaller student models. AED…
-
New distillation method teaches AI models to avoid shortcuts
Researchers have developed a new knowledge distillation technique called Anti-Shortcut Distillation (ASD). This method uses an early-stage teacher model as a negative reference to guide a student model away from learnin…
-
New distillation method improves autoregressive video generation
Researchers have developed Context-Matched Distillation (CMD), a novel framework for improving autoregressive video generation models. CMD aligns the teacher model's supervision with the causal context available to the …
-
Research benchmarks SLM trustworthiness: quantization outperforms pruning
A new research paper explores the trustworthiness of small language models (SLMs) by comparing pre-trained models with compressed versions. The study found that quantization is more effective than network pruning in mai…
-
New SQuaT framework enhances self-supervised knowledge distillation for low-bit models
Researchers have developed SQuaT, a novel framework for self-supervised knowledge distillation that addresses limitations in existing methods when combining quantization-aware training with distillation. SQuaT theoretic…
-
ByteDance founder Zhang Yiming returns, halts AI model distillation
ByteDance founder Zhang Yiming has returned to the company and instructed the Seed AI research team to cease model distillation. He believes this practice, which involves training smaller models on larger ones, quickly …