PulseAugur
实时 10:20:16
English(EN) SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

新的语料库和框架评估大型语言模型写作反馈的质量

研究人员推出了 SEFORA,这是一个旨在捕捉教师如何对学生写作提供反馈的新语料库,以及 UniMatch,一个用于评估人工智能生成反馈质量的评估框架。SEFORA 包含 564 篇学生论文草稿的 8,000 多个教师注释。UniMatch 框架衡量人工智能反馈单元与教师设定的标准之间的语义对应关系和一致性。使用 UniMatch 进行的实验表明,当前的大型语言模型在生成符合教师优先事项的反馈方面存在困难,随着生成反馈的增多,性能会下降,并且没有经过测试的配置能超过 0.4 的 F1 分数。 AI

影响 这项研究突显了当前大型语言模型在写作反馈方面的局限性,表明需要更好地与人类教师的优先事项保持一致。

排序理由 该集群描述了一篇介绍用于人工智能生成写作反馈的语料库和评估框架的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的语料库和框架评估大型语言模型写作反馈的质量

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Shayan Peyghambari Oskoui, Norah Almousa, Zhaoyi Joey Hou, Carolina Gustafson, Gayle Rogers, Raquel Coelho, Diane Litman, Xiang Lorraine Li ·

    SEFORA:带反馈的学生论文语料库和LLM反馈评估框架

    arXiv:2607.00274v1 Announce Type: cross Abstract: Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a natural path to scaling writing support, but two gaps stand in the way: few public corpora c…

  2. arXiv cs.CL TIER_1 English(EN) · Xiang Lorraine Li ·

    SEFORA:带反馈的学生论文语料库和LLM反馈评估框架

    Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a natural path to scaling writing support, but two gaps stand in the way: few public corpora capture how instructors actually deliver feedback i…