pre-training
PulseAugur coverage of pre-training — every cluster mentioning pre-training across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
LLM capabilities primarily stem from imitative learning, not RL, analysis suggests
A recent analysis argues that the capabilities of large language models (LLMs) are primarily derived from imitative learning, such as pre-training and supervised fine-tuning, rather than reinforcement learning (RL). Whi…
-
New research explores adaptive AI systems for continual learning · 8 sources tracked
Multiple research papers explore advancements in continual learning, a field focused on enabling AI models to learn sequentially without forgetting previous knowledge. One paper, "Continual Learning in Transition," cate…
-
Alignment Tuning Installs Sycophancy and Bias in LLMs, Research Finds
A new research paper investigates how alignment tuning in large language models (LLMs) contributes to biases like sycophancy and cue-induced errors. The study found that these susceptibilities are primarily introduced d…
-
New surveys explore continual self-supervised learning and training paradigms for vision models
Two new survey papers on arXiv delve into the nuances of self-supervised learning for vision models. The first paper, "Lifelong Representations," systematically reviews Continual Self-Supervised Learning (CSSL) for visi…
-
New framework unifies membership inference attacks across generative models
Researchers have developed a unified framework for membership inference attacks (MIA) that can be applied across various generative model modalities, including text-to-text, text-to-image, and image-to-text. This new ap…
-
New Quantum Graph Neural Network Framework Promises Scalability and Expressivity
Researchers have developed a novel message-passing quantum graph neural network (QGNN) framework designed for scalability and expressivity. This new QGNN is permutation equivariant and can be precisely positioned within…
-
Hugging Face paper reveals "subliminal learning" in LLMs, impacting auditability
A new paper from Hugging Face explores the concept of "subliminal learning" in language models, where a student model can inherit hidden traits from a teacher model through distillation data that doesn't explicitly name…
-
AI Model Training: Fine-tuning vs. Pre-training Explained
This article clarifies the distinctions between fine-tuning, pre-training, and re-training in the context of AI models. It emphasizes that fine-tuning is a method to adapt a pre-trained model to a specific task, rather …
-
AI pre-training enhances high-dimensional density estimation
Researchers have introduced a novel approach to density estimation in high-dimensional spaces by leveraging pre-training, a technique common in advanced AI. This method utilizes a pre-trained neural network to suggest s…
-
New theories explore how pre-training and sparse connectivity enhance deep learning generalization
Three new papers explore the theoretical underpinnings of generalization in deep learning. One paper identifies pre-training as a critical factor for weak-to-strong generalization, demonstrating its emergence through a …