PulseAugur
实时 10:56:16
English(EN) Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)

新的大语言模型工具评估论文评分偏见和学生AI依赖性

研究人员开发了新的工具和分析方法来评估大语言模型(LLMs)在学术写作中的性能和公平性。一项研究介绍了WrAFT,一个用于自动评分和反馈的模块化系统,该系统使用Llama 3.3 70B Instruct和GPT-4o等模型实现了最先进的性能。另一篇论文调查了针对TOEFL论文自动评分的大语言模型中的第一语言偏见,发现来自欧洲语言背景的论文得分高于来自东亚背景的论文,尽管跨提示泛化稳定。此外,还开发并验证了一个新的量表GenAI-RTS,用于衡量学生在学术写作中对生成式AI的依赖程度,并将依赖性分为战略型、工具型、依赖型和对话型。 AI

影响 这些研究为评估大语言模型的公平性以及理解学生在学术写作中与AI的互动提供了新的工具和见解。

排序理由 多篇学术论文发布在arXiv上,详细介绍了关于大语言模型在自动评分、偏见检测和学生AI依赖性方面的新研究。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 8 个来源。 我们如何撰写摘要 →

新的大语言模型工具评估论文评分偏见和学生AI依赖性

报道来源 [8]

  1. arXiv cs.CL TIER_1 English(EN) · John Maurice Gayed ·

    调查大型语言模型自动作文评分中的母语偏见:一项基于跨提示的开放权重AI模型在托福作文上的评估

    arXiv:2607.14605v1 Announce Type: new Abstract: This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and infere…

  2. arXiv cs.AI TIER_1 English(EN) · Shahin Hossain, Tukhbita Afroz Nawmi ·

    衡量学生在学术写作中对生成式AI的依赖程度:生成式AI依赖类型量表(GenAI-RTS)的开发与多源验证

    arXiv:2607.14301v1 Announce Type: new Abstract: As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than simply whether they use them, has become a central question for learning, academic integrity, an…

  3. arXiv cs.AI TIER_1 English(EN) · Adnan Labib, Yixuan Huang, Jiahui Wu, John Maurice Gayed, Zheng Yuan, Qiao Wang ·

    WrAFT:用于议论文的模块化自动写作评估系统

    arXiv:2607.14524v1 Announce Type: new Abstract: This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays. WrAFT adopts a modular design by dividing automated wri…

  4. arXiv cs.CL TIER_1 English(EN) · Steven Coyne, Diana Galvan-Sosa, Ryan Spring, Machi Shimmei, Michael Zock, Keisuke Sakaguchi, Kentaro Inui ·

    AI生成的反馈效果如何?对超过20,000份EFL论文草稿的内在和外在评估

    arXiv:2607.14591v1 Announce Type: new Abstract: This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best prac…

  5. arXiv cs.CL TIER_1 English(EN) · John Maurice Gayed ·

    调查大型语言模型自动作文评分中的母语偏见:一项针对TOEFL作文的跨提示开放权重AI模型评估

    This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in "AiAWE: An Open-So…

  6. arXiv cs.CL TIER_1 English(EN) · Kentaro Inui ·

    AI 生成的反馈效果如何?对 20,000 多篇 EFL 论文草稿的内在和外在评估

    This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best practices remains an ongoing challenge. WCF meeting …

  7. arXiv cs.CL TIER_1 English(EN) · Qiao Wang ·

    WrAFT:用于议论文的模块化自动写作评估系统

    This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays. WrAFT adopts a modular design by dividing automated writing evaluation (AWE) tasks into scoring, surfac…

  8. arXiv cs.CL TIER_1 English(EN) · Tukhbita Afroz Nawmi ·

    衡量学生在学术写作中对生成式AI的依赖程度:生成式AI依赖类型量表(GenAI-RTS)的开发与多源验证

    As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than simply whether they use them, has become a central question for learning, academic integrity, and educational equity. Existing measures of relia…