PulseAugur
EN
LIVE 06:09:13
ENTITY supervised fine-tuning

supervised fine-tuning

PulseAugur coverage of supervised fine-tuning — every cluster mentioning supervised fine-tuning across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
56
164 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
52
150 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

23 day(s) with sentiment data

RECENT · PAGE 1/9 · 164 TOTAL
  1. TOOL · CL_193541 ·

    New VIGIL system uses LMMs for precise visual distortion detection

    Researchers have introduced VIGIL, a new system designed for precise visual distortion detection in user-generated images. Unlike previous methods that rely on text-driven supervised fine-tuning of large multimodal mode…

  2. TOOL · CL_193330 ·

    TRACE-Memory framework enhances personalized generation by selectively using user history

    Researchers have developed TRACE-Memory, a novel two-stage framework designed to enhance personalized generation systems. This framework selectively incorporates user history only when it provides utility beyond publicl…

  3. RESEARCH · CL_191208 ·

    New agentic approaches enhance knowledge graph question answering and generation

    Researchers are developing agentic systems to improve question answering over knowledge graphs. One approach, "Researcher Agents," focuses on self-improvement by iteratively testing and modifying its own prompts and cod…

  4. TOOL · CL_189503 ·

    LLM Deception Reduced by Self-Other Overlap Training

    Researchers have found that supervised fine-tuning (SFT) can significantly reduce deception in large language models by inducing self-other overlap. Models like Qwen2.5-14B-Instruct, Gemma-3-27B-It, Qwen2.5-32B-Instruct…

  5. RESEARCH · CL_191188 ·

    New framework aligns recommender foundation models with business metrics · 2 sources tracked

    Researchers have developed a novel three-phase post-training framework to better align recommender foundation models with business metrics. This progressive approach separates downstream adaptation, using Linear Probing…

  6. TOOL · CL_187392 ·

    AuroSFT framework improves multi-task fine-tuning by managing adapter states

    Researchers have introduced AuroSFT, a new framework for multi-task fine-tuning that efficiently manages adapter states instead of full model checkpoints. This approach freezes the pretrained backbone and trains only in…

  7. TOOL · CL_185289 ·

    New RL strategy trains MLLMs to refuse non-existent objects

    Researchers have developed a new reinforcement learning strategy called Refusal-Calibrated Group Relative Policy Optimization (RC-GRPO) to improve the ability of Multimodal Large Language Models (MLLMs) to correctly ide…

  8. TOOL · CL_184554 ·

    Fireworks AI expands event series and adds DeepSeek V4 Flash 0731 fine-tuning

    Fireworks AI is expanding its offerings by announcing a new edition of "The AI Dev Stack" event in San Francisco, following its initial event in New York City. This event, scheduled for August 18th, will feature discuss…

  9. RESEARCH · CL_184909 ·

    ToolArtist model integrates reasoning, tool use, and image generation

    Researchers have introduced ToolArtist, a novel agentic image generation model designed to overcome limitations in current text-to-image systems. Unlike previous models with fixed workflows or partial agent control, Too…

  10. TOOL · CL_181201 ·

    Hugging Face's TRL library scores mid-tier on Olud Pulse alignment benchmark

    Hugging Face's Transformer Reinforcement Learning (TRL) library has received a score of 51/100 on the Olud Pulse benchmark. This score places TRL in the middle tier of alignment tools, indicating it is a functional, tho…

  11. TOOL · CL_180596 ·

    Research links learning rates to catastrophic overtraining in LLMs

    A new research paper explores how learning rates impact catastrophic overtraining during the supervised fine-tuning (SFT) of large language models (LLMs). The study, published on arXiv, suggests that different learning …

  12. TOOL · CL_178378 ·

    New GeoRA method enhances RLVR for large language models

    Researchers have introduced GeoRA, a novel low-rank adaptation method specifically designed for Reinforcement Learning with Verifiable Rewards (RLVR). Unlike existing methods that focus on supervised fine-tuning, GeoRA …

  13. TOOL · CL_182288 ·

    Wnuan pipeline enhances enterprise QA models while retaining general capabilities

    Researchers have introduced Wnuan, a novel three-stage pipeline designed to equip models with proprietary enterprise knowledge without sacrificing general capabilities. This method involves constructing task-specific su…

  14. TOOL · CL_176434 ·

    New tool converts agent failures into fine-tuning data

    A new open-source tool called trace2train has been released to convert failed agent traces into supervised fine-tuning (SFT) or Direct Preference Optimization (DPO) training data. Developed as a local CLI tool, it aims …

  15. RESEARCH · CL_180478 ·

    New method RSTG improves LLM reinforcement learning with adaptive teacher guidance

    Researchers have developed RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), a novel method to improve reinforcement learning for large language models. Existing methods like GRPO struggle with sparse re…

  16. TOOL · CL_174181 ·

    New HARGO method optimizes LLMs for diverse HPC tasks

    Researchers have developed HARGO, a novel optimization technique designed to improve the performance of large language models (LLMs) on diverse high-performance computing (HPC) tasks. Traditional reinforcement learning …

  17. TOOL · CL_178586 ·

    New LLM framework Think2Go enhances POI recommendations

    Researchers have introduced Think2Go, a new generative recommendation framework designed to improve next Point-of-Interest (POI) suggestions. This framework addresses limitations in existing methods by enhancing the com…

  18. RESEARCH · CL_175944 ·

    New credit assignment methods enhance AI search agent training · 3 sources tracked

    Researchers have developed new methods for training long-horizon search agents, which are AI systems designed to perform complex, multi-step tasks. One approach, ABSeeker, uses Answer-Backtracked Credit Assignment (ABC)…

  19. TOOL · CL_172055 ·

    OmniAD framework enhances industrial anomaly detection with multimodal reasoning

    Researchers have developed OmniAD, a new multimodal reasoning framework designed to detect and analyze industrial anomalies. This system integrates visual and textual reasoning, using a 'Text-as-Mask Encoding' approach …

  20. TOOL · CL_171872 ·

    AI alignment, model organisms, and toy models share SFT lessons

    Researchers have identified shared lessons in supervised fine-tuning (SFT) that can be applied across seemingly disparate fields like AI alignment, model organisms, and toy models. The study demonstrates how techniques …