PulseAugur
EN
LIVE 15:44:41

DeepSeek-V4-Flash-0731 evaluation reveals high cost and low efficiency compared to Kimi K3

A user on Reddit shared their experience evaluating the new DeepSeek-V4-Flash-0731 model, finding it surprisingly inefficient and costly. In a oneshot evaluation across 34 prompts, the model incurred a cost of $1.29 and achieved a score of 2.7/5, while the Kimi K3 model cost only $0.44 for fewer tokens and a higher score of 3.2/5. The user questioned if they were misusing the model or if their experience was typical. AI

IMPACT Highlights potential inefficiencies in newer models, suggesting users should carefully evaluate cost-performance ratios.

RANK_REASON User-generated evaluation and opinion on a model release, not an official announcement.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

DeepSeek-V4-Flash-0731 evaluation reveals high cost and low efficiency compared to Kimi K3

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/kms_dev ·

    DeepSeek-V4-Flash-0731: Oneshot evals, surprisingly not token efficient??

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vcl502/deepseekv4flash0731_oneshot_evals_surprisingly/"> <img alt="DeepSeek-V4-Flash-0731: Oneshot evals, surprisingly not token efficient??" src="https://preview.redd.it/shxbtioh0rgh1.png?width=140&amp;heigh…