PulseAugur
EN
LIVE 13:08:25
ENTITY supervised fine-tuning

supervised fine-tuning

PulseAugur coverage of supervised fine-tuning — every cluster mentioning supervised fine-tuning across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
46
175 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
43
162 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

15 day(s) with sentiment data

RECENT · PAGE 1/10 · 200 TOTAL
  1. TOOL · CL_259290 ·

    New method refines AI agent trajectories, cutting costs and boosting accuracy

    Researchers have developed a method called Dependency-Aware Trajectory Refinement (DATR) to optimize the fine-tuning of multi-turn AI agents. This technique involves representing agent trajectories as a Directed Acyclic…

  2. TOOL · CL_257041 ·

    Pruning LLMs for smart homes: MoE models more resilient than dense

    A new research paper explores the impact of pruning on large language models (LLMs) specifically within the context of smart-home tool calling. The study systematically evaluated pruning-induced degradation across vario…

  3. TOOL · CL_256874 ·

    New framework MOCC-R1 improves multimodal counselor response consistency

    Researchers have introduced MOCC-R1, a novel framework designed to enhance the consistency between reasoning and response generation in multimodal counselor systems. This framework addresses limitations in existing data…

  4. TOOL · CL_256853 ·

    New ReDraft method improves LLM post-training by revising model failures

    Researchers have developed a new method called ReDraft for continually post-training large multimodal models. This technique aims to enhance new capabilities without sacrificing existing ones, a common challenge in mode…

  5. RESEARCH · CL_254525 ·

    New research explores data synthesis and curriculum learning for advanced LLM training

    Two new research papers explore advanced techniques for training large language models (LLMs) using Reinforcement Learning with Verifiable Rewards (RLVR). The first paper introduces MIFS, a pipeline for synthesizing RL-…

  6. TOOL · CL_253866 ·

    Constitutional AI: Principles-Based LLM Alignment Explained

    Constitutional AI (CAI) offers a novel approach to aligning large language models (LLMs) by using a set of predefined principles, or a "constitution," rather than relying solely on human feedback. This method involves a…

  7. TOOL · CL_252031 ·

    New testbed reveals reward hacking in LLMs emerges during fine-tuning

    Researchers have introduced Countdown-Code, a new testbed designed to accurately measure reward hacking in large language models. This environment separates true task rewards from proxy rewards, revealing that reward ha…

  8. RESEARCH · CL_252071 ·

    MInTRL enhances on-policy RL with minimal interventions · 2 sources tracked

    Researchers have introduced Minimal Intervention Reinforcement Learning (MInTRL), a novel approach to enhance on-policy reinforcement learning. MInTRL expands exploration by incorporating sparse, local interventions int…

  9. RESEARCH · CL_247882 ·

    New benchmark EgoGenEval reveals AI visual generators struggle with physical consistency

    Researchers have introduced EgoGenEval, a new benchmark designed to assess the physical consistency of visual generators, particularly under ego-motion. The benchmark evaluates two key aspects: Camera Motion Grounding (…

  10. TOOL · CL_245380 ·

    New framework tackles adaptive prompt injection attacks on AI agents

    Researchers have developed CoRL, a novel framework for defending against and simulating adaptive indirect prompt injection (IPI) attacks on tool-augmented language agents. IPI attacks hide adversarial instructions withi…

  11. RESEARCH · CL_245203 ·

    New methods enhance LLM financial reasoning without weight modification · 2 sources tracked

    Researchers have developed new methods for adapting large language models (LLMs) to specialized financial reasoning tasks. One approach, ASDA, automatically generates structured skill artifacts without modifying model w…

  12. TOOL · CL_245192 ·

    New GIFT method enhances Large Reasoning Model training by reconciling SFT and RL

    Researchers have introduced GIFT (Gibbs Initialization with Finite Temperature), a novel method to improve the post-training process for Large Reasoning Models (LRMs). This technique addresses the optimization mismatch …

  13. TOOL · CL_242941 ·

    EAGER framework enhances e-commerce query recommendations using LLMs

    Researchers have developed EAGER, a novel two-stage framework designed to generate more relevant query suggestions for e-commerce search. The first stage, enrichment, uses supervised fine-tuning with a curriculum that p…

  14. TOOL · CL_239398 ·

    New LLM predicts pension enrollment in China using policy cues

    Researchers have developed FlexPension-LLM, a specialized large language model designed to predict pension enrollment among flexible workers in China. This model integrates policy-grounded cues, such as marginal effects…

  15. TOOL · CL_239296 ·

    New PetQA benchmark evaluates AI veterinary knowledge

    Researchers have developed PetQA, a new benchmark designed to evaluate the veterinary knowledge and clinical reasoning capabilities of large language models (LLMs) and large vision-language models (LVLMs). The benchmark…

  16. TOOL · CL_239254 ·

    New paper unifies LLM training methods via Bayesian lens

    A new paper proposes a unified Bayesian framework to understand various large language model training and evaluation paradigms, including supervised fine-tuning (SFT), in-context learning (ICL), and KL-regularized reinf…

  17. RESEARCH · CL_239613 ·

    New RAG frameworks enhance multimodal AI for specialized tasks · 2 sources tracked

    Two new research papers explore advancements in multimodal retrieval-augmented generation (RAG) for specialized applications. The first paper introduces a generator-in-the-loop alignment framework to improve the utility…

  18. TOOL · CL_235513 ·

    New SWIM task uses AI to simulate student writing proficiency

    Researchers have introduced SWIM, a new task designed to simulate student writing through proficiency-conditioned generation. The study explored prompting, supervised fine-tuning (SFT), and reinforcement learning (RL) m…

  19. TOOL · CL_235485 ·

    Review details LLM techniques for medical reasoning and future challenges

    A recent systematic review published on arXiv examines the advancements and challenges of large language models (LLMs) in medical reasoning. The paper categorizes techniques for enhancing LLM reasoning into training-tim…

  20. RESEARCH · CL_235470 ·

    New research paper details OPD-then-RL for enhanced LLM reasoning

    A new research paper proposes a two-stage approach called OPD-then-RL for improving large language models' reasoning capabilities. This method combines On-Policy Distillation (OPD) with Reinforcement Learning with Verif…