supervised fine-tuning
PulseAugur coverage of supervised fine-tuning — every cluster mentioning supervised fine-tuning across labs, papers, and developer communities, ranked by signal.
- instance of Direct Preference Optimization 90%
- developed by Grpo 70%
- used by Direct Preference Optimization: Your Language Model is Secretly a Reward Model 70%
- instance of Gotit.pub 70%
- used by alphaXiv 70%
- instance of alphaXiv 70%
- used by Group Relative Policy Optimization 70%
- used by On-Policy Distillation 70%
- used by Qwen3 1.7B 70%
- used by Grpo 60%
- other Direct Preference Optimization 60%
- competes with Direct Preference Optimization 60%
23 day(s) with sentiment data
-
New VIGIL system uses LMMs for precise visual distortion detection
Researchers have introduced VIGIL, a new system designed for precise visual distortion detection in user-generated images. Unlike previous methods that rely on text-driven supervised fine-tuning of large multimodal mode…
-
TRACE-Memory framework enhances personalized generation by selectively using user history
Researchers have developed TRACE-Memory, a novel two-stage framework designed to enhance personalized generation systems. This framework selectively incorporates user history only when it provides utility beyond publicl…
-
New agentic approaches enhance knowledge graph question answering and generation
Researchers are developing agentic systems to improve question answering over knowledge graphs. One approach, "Researcher Agents," focuses on self-improvement by iteratively testing and modifying its own prompts and cod…
-
LLM Deception Reduced by Self-Other Overlap Training
Researchers have found that supervised fine-tuning (SFT) can significantly reduce deception in large language models by inducing self-other overlap. Models like Qwen2.5-14B-Instruct, Gemma-3-27B-It, Qwen2.5-32B-Instruct…
-
New framework aligns recommender foundation models with business metrics · 2 sources tracked
Researchers have developed a novel three-phase post-training framework to better align recommender foundation models with business metrics. This progressive approach separates downstream adaptation, using Linear Probing…
-
AuroSFT framework improves multi-task fine-tuning by managing adapter states
Researchers have introduced AuroSFT, a new framework for multi-task fine-tuning that efficiently manages adapter states instead of full model checkpoints. This approach freezes the pretrained backbone and trains only in…
-
New RL strategy trains MLLMs to refuse non-existent objects
Researchers have developed a new reinforcement learning strategy called Refusal-Calibrated Group Relative Policy Optimization (RC-GRPO) to improve the ability of Multimodal Large Language Models (MLLMs) to correctly ide…
-
Fireworks AI expands event series and adds DeepSeek V4 Flash 0731 fine-tuning
Fireworks AI is expanding its offerings by announcing a new edition of "The AI Dev Stack" event in San Francisco, following its initial event in New York City. This event, scheduled for August 18th, will feature discuss…
-
ToolArtist model integrates reasoning, tool use, and image generation
Researchers have introduced ToolArtist, a novel agentic image generation model designed to overcome limitations in current text-to-image systems. Unlike previous models with fixed workflows or partial agent control, Too…
-
Hugging Face's TRL library scores mid-tier on Olud Pulse alignment benchmark
Hugging Face's Transformer Reinforcement Learning (TRL) library has received a score of 51/100 on the Olud Pulse benchmark. This score places TRL in the middle tier of alignment tools, indicating it is a functional, tho…
-
Research links learning rates to catastrophic overtraining in LLMs
A new research paper explores how learning rates impact catastrophic overtraining during the supervised fine-tuning (SFT) of large language models (LLMs). The study, published on arXiv, suggests that different learning …
-
New GeoRA method enhances RLVR for large language models
Researchers have introduced GeoRA, a novel low-rank adaptation method specifically designed for Reinforcement Learning with Verifiable Rewards (RLVR). Unlike existing methods that focus on supervised fine-tuning, GeoRA …
-
Wnuan pipeline enhances enterprise QA models while retaining general capabilities
Researchers have introduced Wnuan, a novel three-stage pipeline designed to equip models with proprietary enterprise knowledge without sacrificing general capabilities. This method involves constructing task-specific su…
-
New tool converts agent failures into fine-tuning data
A new open-source tool called trace2train has been released to convert failed agent traces into supervised fine-tuning (SFT) or Direct Preference Optimization (DPO) training data. Developed as a local CLI tool, it aims …
-
New method RSTG improves LLM reinforcement learning with adaptive teacher guidance
Researchers have developed RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), a novel method to improve reinforcement learning for large language models. Existing methods like GRPO struggle with sparse re…
-
New HARGO method optimizes LLMs for diverse HPC tasks
Researchers have developed HARGO, a novel optimization technique designed to improve the performance of large language models (LLMs) on diverse high-performance computing (HPC) tasks. Traditional reinforcement learning …
-
New LLM framework Think2Go enhances POI recommendations
Researchers have introduced Think2Go, a new generative recommendation framework designed to improve next Point-of-Interest (POI) suggestions. This framework addresses limitations in existing methods by enhancing the com…
-
New credit assignment methods enhance AI search agent training · 3 sources tracked
Researchers have developed new methods for training long-horizon search agents, which are AI systems designed to perform complex, multi-step tasks. One approach, ABSeeker, uses Answer-Backtracked Credit Assignment (ABC)…
-
OmniAD framework enhances industrial anomaly detection with multimodal reasoning
Researchers have developed OmniAD, a new multimodal reasoning framework designed to detect and analyze industrial anomalies. This system integrates visual and textual reasoning, using a 'Text-as-Mask Encoding' approach …
-
AI alignment, model organisms, and toy models share SFT lessons
Researchers have identified shared lessons in supervised fine-tuning (SFT) that can be applied across seemingly disparate fields like AI alignment, model organisms, and toy models. The study demonstrates how techniques …