PulseAugur
实时 17:56:50
English(EN) Human Grounded Evaluation of Large Language Models for Optical Network Automation

新的评估流程用于评估语言大模型在光网络自动化中的应用

研究人员开发了HuGLEN,一个新颖的评估流程,旨在评估语言大模型(LLMs)在光网络自动化中的应用。该流程采用“以LLM为裁判”的方法,并结合专家评分,创建了一种可扩展且可复现的比较方法。该系统旨在根据质量效率得分(QES)对LLMs进行排名,该得分平衡了解释质量和推理成本。结果表明,一个12B参数的LLM取得了最高的QES,证明了其在面向操作员的自动化任务中的最佳权衡。 AI

影响 这一新的评估框架可以简化网络自动化任务中LLMs的选择,提高效率和质量。

排序理由 该集群描述了一篇研究论文,详细介绍了一个用于特定领域LLMs的新评估流程。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的评估流程用于评估语言大模型在光网络自动化中的应用

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Kiarash Rezaei, Omran Ayoub, Paolo Monti, Carlos Natalino ·

    面向光网络自动化的面向人类的大型语言模型评估

    arXiv:2607.18068v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an L…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Human Grounded Evaluation of Large Language Models for Optical Network Automation

    Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an LLM-as-a-judge together with a small set of expert …