PulseAugur
实时 10:12:13
实体 Reinforcement Learning with AI Feedback

Reinforcement Learning with AI Feedback

PulseAugur coverage of Reinforcement Learning with AI Feedback — every cluster mentioning Reinforcement Learning with AI Feedback across labs, papers, and developer communities, ranked by signal.

Show in brief
总计 · 30天
1
90 天内 1
发布 · 30天
0
90 天内 0
论文 · 30天
1
90 天内 1
层级分布 · 90 天
主题
情绪 · 30 天

1 天有情绪数据

最近 · 第 1/1 页 · 共 1 条
  1. TOOL · CL_180484 ·

    新型波斯语医学大语言模型家族“Gaokerena”发布

    研究人员推出 Gaokerena,一个专为消费级硬件设计的新型小型波斯语医学语言模型家族。该模型家族包括 Gaokerena-V,它在一个包含 9000 万个 token 的波斯语医学语料库上进行了训练;以及 Gaokerena-R,它集成了思维链(Chain-of-Thought)和 RLAIF 以提高临床推理能力。尽管 Gaokerena-R 使用的数据集较小,但其基准测试得分高于 Gaokerena-V。两个模型都配备了不确定性…