PulseAugur
实时 16:34:26
English(EN) DeepSeek-V4-Flash-0731: Oneshot evals, surprisingly not token efficient??

DeepSeek-V4-Flash-0731 评估显示,与 Kimi K3 相比,成本高且效率低

一位 Reddit 用户分享了他们对新 DeepSeek-V4-Flash-0731 模型进行评估的经验,发现其效率低下且成本高昂。在 34 个提示的单次评估中,该模型的成本为 1.29 美元,得分为 2.7/5,而 Kimi K3 模型仅花费 0.44 美元,使用更少的代币,得分却高达 3.2/5。该用户质疑是自己误用了模型,还是他们的经历具有代表性。 AI

影响 凸显了新模型中潜在的低效率,建议用户仔细评估成本效益比。

排序理由 用户生成的模型发布评估和意见,并非官方公告。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

DeepSeek-V4-Flash-0731 评估显示,与 Kimi K3 相比,成本高且效率低

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/kms_dev ·

    DeepSeek-V4-Flash-0731: Oneshot evals, surprisingly not token efficient??

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vcl502/deepseekv4flash0731_oneshot_evals_surprisingly/"> <img alt="DeepSeek-V4-Flash-0731: Oneshot evals, surprisingly not token efficient??" src="https://preview.redd.it/shxbtioh0rgh1.png?width=140&amp;heigh…