PulseAugur
中
实时 11:28:11
English(EN) Write the eval before the prompt

LLM开发:先写评估,再写提示词,以获得可靠的功能

通过采用类似传统软件工程的测试驱动方法,可以改进LLM功能的开发。这包括在编写提示词之前,创建一个包含预期结果的“评估”数据集。这个评估数据集应该基于实际的失败和纠正来构建,而不是凭空想象的场景,以准确反映生产中的问题。通过将提示词应用于这个全面的评估数据集,开发人员可以识别回归并可靠地提高LLM的性能。 AI

影响 采用具有全面评估数据集的测试驱动方法,可以带来更可靠、更健壮的LLM功能开发。

排序理由 该条目讨论的是LLM开发的一种方法论,而不是一个特定的发布或事件。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLM开发:先写评估,再写提示词,以获得可靠的功能

本文如何被排名

Signal score
4 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
该条目讨论的是LLM开发的一种方法论,而不是一个特定的发布或事件。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Ahmet Zeybek ·

    先写评估,再写提示

    <p>This is how I used to build an LLM feature. Write a prompt. Try it on the four or five inputs I had to hand, and adjust it until those looked right. Ship it behind a flag. Wait for the first bug report, which came within a day and was about an input nothing like my five. Adjus…