PulseAugur
EN
LIVE 11:09:54
中文(ZH) DeepSeek低价风暴打服硅谷!海外平台争相倒贴V4 Flash

DeepSeek V4 Flash challenges top AI models with low-cost, high-performance release · 10 sources tracked

DeepSeek has released its V4 Flash model, which offers performance comparable to top-tier models like OpenAI's GPT-5.6 Luna and Anthropic's Claude Opus 4.8, but at a significantly lower cost. This new model, particularly the 0731 update, has seen substantial gains in agentic and coding benchmarks due to re-post-training, making it a highly attractive option for developers and businesses seeking cost-effective AI solutions. The V4 Flash's affordability and strong performance are reshaping developer defaults and are seen as a pivotal moment for Chinese AI development, potentially democratizing access to advanced AI capabilities. AI

IMPACT Accelerates adoption of advanced AI by making high-performance models accessible and affordable for developers and businesses.

RANK_REASON Frontier-lab model release with performance and pricing details.

Read on 量子位 (QbitAI) →

AI-generated summary · Google Gemini · from 110 sources. How we write summaries →

DeepSeek V4 Flash challenges top AI models with low-cost, high-performance release · 10 sources tracked

COVERAGE [110]

  1. Unsloth — Releases TIER_1 English(EN) · danielhanchen ·

    Faster downloading + DeepSeek-V4 Flash 0731

    <p>Hey everyone! For folks who missed the news - Kimi K3 &amp; DeepSeek v4 Flash can now run locally with Unsloth Dynamic GGUFs! We now added more efficient and faster downloading for Colab, low memory systems and also high memory and CPU systems - we auto fallback to HTTP as wel…

  2. Unsloth — Releases TIER_1 English(EN) · danielhanchen ·

    Faster downloading + DeepSeek-V4 Flash 0731

    <p>Hey everyone! For folks who missed the news - Kimi K3 &amp; DeepSeek v4 Flash can now run locally with Unsloth Dynamic GGUFs! We now added more efficient and faster downloading for Colab, low memory systems and also high memory and CPU systems - we auto fallback to HTTP as wel…

  3. Unsloth — Releases TIER_1 English(EN) · danielhanchen ·

    Faster downloading + DeepSeek-V4 Flash 0731

    <p>Hey everyone! For folks who missed the news - Kimi K3 &amp; DeepSeek v4 Flash can now run locally with Unsloth Dynamic GGUFs! We now added more efficient and faster downloading for Colab, low memory systems and also high memory and CPU systems - we auto fallback to HTTP as wel…

  4. 量子位 (QbitAI) TIER_1 中文(ZH) · Jay ·

    DeepSeek's low-price storm conquers Silicon Valley! Overseas platforms rush to subsidize V4 Flash

    那还说啥了梁圣,我订阅费全给你就是了呗

  5. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI.

    DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO, then deploy the fine-tuned model on Together AI for production inference. https://t.co/MFkqLlOKAj https://t.co/TpTihq8vic

  6. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance available. Open weights,

    Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance available. Open weights, zero data retention, U.S. &amp; EU hosting.

  7. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Congrats to @deepseek_ai on their release of DeepSeekv4 Flash 0731 🔥 It massively beats Nemotron3 Ultra on agentic tasks while having 4.2x fewer active paramete

    Congrats to @deepseek_ai on their release of DeepSeekv4 Flash 0731 🔥 It massively beats Nemotron3 Ultra on agentic tasks while having 4.2x fewer active parameters and close to 2x fewer total parameters! Committee-based model frontier development does not work. A focused team is …

  8. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    DeepSeek V4 Flash 0728 really is the current cost-performance frontier!

    DeepSeek V4 Flash 0728 really is the current cost-performance frontier!

  9. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE.

    We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna’s performance at roughly 1/6 the cost. More insights in the thread! 👇

  10. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    🤯 DeepSeek reports V4 Flash at 82.7 on Terminal-Bench 2.1, ahead of V4 Pro Preview at 72.1, despite using roughly one-fifth the total parameters. https://t.co/S

    🤯 DeepSeek reports V4 Flash at 82.7 on Terminal-Bench 2.1, ahead of V4 Pro Preview at 72.1, despite using roughly one-fifth the total parameters. https://t.co/SGRiiKCcu3

  11. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    RT @zainhas: DeepSeek-V4 Flash vs. GPT 5.6 Luna is crazy! https://t.co/HhTiVHnbgf

    RT @zainhas: DeepSeek-V4 Flash vs. GPT 5.6 Luna is crazy! https://t.co/HhTiVHnbgf

  12. 36氪 (36Kr) TIER_1 中文(ZH) ·

    DeepSeek-V4-Flash Official Version API Lands on National Supercomputing Internet

    36氪获悉,近日,DeepSeek-V4-Flash正式版API面向公众开启公测,国家超算互联网第一时间同步上线DeepSeek-V4-Flash模型API调用与下载服务。

  13. TLDR AI TIER_1 English(EN) · TLDR ·

    DeepSeek V4 Flash ⚡, OpenAI’s math breakthrough 🔢, Qwen 3.8-Max 🤖

  14. 36氪 (36Kr) TIER_1 中文(ZH) ·

    DeepSeek: DeepSeek-V4-Pro Official Version Will Be Released Soon

    36氪获悉,DeepSeek发布更新日志,DeepSeek-V4-Flash正式版API上线公测。DeepSeek-V4-Flash-0731的模型结构、尺寸和DeepSeek-V4-Flash-preview保持一致,仅重新进行了后训练。DeepSeek-V4-Pro正式版将会尽快发布。

  15. The Decoder TIER_1 English(EN) · Thomas Joos ·

    New Deepseek Flash model matches OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost

    <p><img alt="" class="attachment-full size-full wp-post-image" height="1152" src="https://the-decoder.com/wp-content/uploads/2026/06/deepseek_red_whale.png" style="height: auto; margin-bottom: 10px;" width="2048" /></p> <p> Deepseek's budget model V4 Flash gets a major boost with…

  16. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    DeepSeek V4-Flash Tops Global Token Consumption: Chinese Models Sweep OpenRouter Top 5 as Price-Performance Kill Line Reshapes Developer Defaults

    DeepSeek V4-Flash leads OpenRouter weekly rankings with 7.1 trillion tokens, Chinese models claim nine of the global top 10 slots, and global weekly AI token usage crosses 56.8 trillion.

  17. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    DeepSeek-V4-Flash Official Version Released and Open-Sourced: 304B Lightweight Model Surpasses V4-Pro Preview and Matches Claude Opus-4.8, Hailed as Third DeepSeek Moment

    DeepSeek-V4-Flash-0731 official release outperforms V4-Pro preview with 82.7 performance score at 0.14/0.28 USD pricing, tops VulcanBench, and reaches HuggingFace trending second.

  18. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains

    <p>DeepSeek published DeepSeek-V4-Flash-0731 on Hugging Face and moved the official V4-Flash API into public beta on July 31, 2026. The model card is explicit that this is the official release superseding the preview, and that the architecture and size are unchanged. The gains co…

  19. Towards AI TIER_1 English(EN) · Caspar Bannink - AI Engineer ·

    Deepseek Did It Again, V4 Flash GA Is The Best Model Per Dollar

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/deepseek-did-it-again-v4-flash-ga-is-the-best-model-per-dollar-cae8661d962e?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1400/1*hgYzGY5jaXpgnUMfoZYVNw.jp…

  20. Towards AI TIER_1 English(EN) · allglenn ·

    DeepSeek V4 Flash 0731 Outscores V4 Pro at 5x Lower Cost

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/deepseek-v4-flash-0731-outscores-v4-pro-at-5x-lower-cost-cd81d817e982?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1536/1*d2asC8IzuphoDaVbG9QEmw.png" wid…

  21. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    A 304B parameter open-source model just matched Claude Opus-4.8. Beijing-based DeepSeek’s official release of V4-Flash marks a massive leap in Chinese AI effici

    A 304B parameter open-source model just matched Claude Opus-4.8. Beijing-based DeepSeek’s official release of V4-Flash marks a massive leap in Chinese AI efficiency, outperforming its own V4-Pro preview. This "third DeepSeek moment" gives global developers a formidable, highly ef…

  22. Medium — Claude tag TIER_1 English(EN) · Joe Njenga ·

    I Tried DeepSeek V4-Flash on Claude Code (It Beats Fable 5 at 71x Lower Cost)

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@joe.njenga/i-tried-deepseek-v4-flash-on-claude-code-it-beats-fable-5-at-71x-lower-cost-fa5b74669412?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1280/1*mklI1uM1G6qwV…

  23. Mastodon — sigmoid.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @u1tra_instinct: 🚨🚨🚨🚨🚨: most requested since the DeepSeek-V4 Flash GA release on 07/31. Yesterday. Now obliterated 32/32, 100% compatible m

    RT @u1tra_instinct: 🚨🚨🚨🚨🚨: am meisten angefragt seit der DeepSeekV4-Flash-GA-Veröffentlichung vom 31.07. Gestern. Jetzt abliteriert 32/32, zu 100 % kompatibel mit DSpark. Wenn euch meine Arbeit gefällt und ihr beitragen möchtet, könnt ihr mir gerne über X-Money oder GoFundMe (Lin…

  24. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    DeepSeek V4 Flash 0731 is interesting not because it became larger, but because it became much more capable through post-training. The same architecture now del

    DeepSeek V4 Flash 0731 is interesting not because it became larger, but because it became much more capable through post-training. The same architecture now delivers stronger coding-agent and tool-use performance while remaining remarkably inexpensive through the API. https:// op…

  25. Medium — Claude tag TIER_1 English(EN) · Chimin ·

    DeepSeek-V4-Flash Public Beta: Claude Code and Codex Can Both Be Used at Dirt-Cheap Prices…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://githubdaily.medium.com/deepseek-v4-flash-public-beta-claude-code-and-codex-can-both-be-used-at-dirt-cheap-prices-29a955104c96?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1742/1…

  26. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    https://www. europesays.com/3166986/ DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains # AgenticAI # AgenticArtificialIntelligence #

    https://www. europesays.com/3166986/ DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIntelligence

  27. Medium — Claude tag TIER_1 English(EN) · inprogrammer ·

    DeepSeek V4 Flash Is 30x Cheaper Than Claude. I Still Switched Back.

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/data-science-collective/deepseek-v4-flash-is-30x-cheaper-than-claude-i-still-switched-back-2ac429bf4e14?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1024/1*ZKNqc_ARJ4…

  28. Towards AI TIER_1 Nederlands(NL) · Mia Efoxtech ·

    DeepSeek V4 vs DeepSeek V4 Flash: Which Model Should Developers Choose in 2026?

    <p>Choose DeepSeek V4 Flash for high-volume, latency-sensitive, and cost-controlled workloads; choose DeepSeek V4 Pro for difficult reasoning, coding, research, and long-horizon agents. Both offer a 1M-token context window, 384K maximum output, thinking and non-thinking modes, JS…

  29. r/LocalLLaMA TIER_1 English(EN) · /u/stereohype ·

    DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide

    <!-- SC_OFF --><div class="md"><p>Been benchmarking DSv4 Flash 0731 on a Flow Z13 (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151, 128GB LPDDR5X) for the past week. Figured I'd share what actually works and what doesn't — there are a lot of gotchas on this hardware.</p> <h2>Results</h…

  30. r/LocalLLaMA TIER_1 English(EN) · /u/Striking-Swim6702 ·

    DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integration

    <!-- SC_OFF --><div class="md"><p>Spent two nights getting <code>deepseek-ai/DeepSeek-V4-Flash-0731</code> (284B MoE, 13B active, native FP4/FP8, 1M context) running production-grade on two DGX Sparks connected by one QSFP DAC cable. Everything — scripts, tuning data, raw benchma…

  31. r/LocalLLaMA TIER_1 English(EN) · /u/Porespellar ·

    DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks

    <!-- SC_OFF --><div class="md"><p>Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getting a lot of people to buy a couple of NVIDIA G…

  32. r/LocalLLaMA TIER_1 English(EN) · /u/Exciting-Camera3226 ·

    DeepSeek V4 Flash 0731 hits 82.7% on Terminal-Bench 2.1 in an independent public-harness run (445 trials)

    <!-- SC_OFF --><div class="md"><p>Disclosure: I’m the author of Ante.</p> <p>DeepSeek recently reported an 82.7% score on Terminal-Bench 2.1 for DeepSeek V4 Flash 0731. Its evaluation used “DeepSeek Harness minimal mode,” which hasn’t been released yet.</p> <p>We wanted to see wh…

  33. r/LocalLLaMA TIER_1 English(EN) · /u/corruptbytes ·

    Updated benchmark: Deepseek V4 Flash on SlopCodeBench (local)

    <!-- SC_OFF --><div class="md"><p>Howdy - I posted a benchmark here - <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vbtiy7/deepseek_v4_flash_on_slopcodebench/">https://www.reddit.com/r/LocalLLaMA/comments/1vbtiy7/deepseek_v4_flash_on_slopcodebench/</a></p> <p>This was us…

  34. r/LocalLLaMA TIER_1 English(EN) · /u/nomorebuttsplz ·

    any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on?

    <!-- SC_OFF --><div class="md"><p>I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 million tokens?</p> </div><!-- SC_ON --> &#32; …

  35. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek V4 Flash just jumped 56% in price as SGLang adds Kimi K3 support, two signals from the same hour that can change your deployment math. https:// olud.ai

    DeepSeek V4 Flash just jumped 56% in price as SGLang adds Kimi K3 support, two signals from the same hour that can change your deployment math. https:// olud.ai/news/2026-08-08.html # AI # OpenSource # AINews

  36. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek V4 Flash closes in on frontier benchmarks at bargain prices, OpenAI pauses its Astra model over autonomous hacking fears, and memory chip shortages now

    DeepSeek V4 Flash closes in on frontier benchmarks at bargain prices, OpenAI pauses its Astra model over autonomous hacking fears, and memory chip shortages now extend through 2027. https:// ai0.news/posts/2026-08-08-dail y-digest/ # AI # Cybersecurity # OpenAI # AiPolicy

  37. r/LocalLLaMA TIER_1 English(EN) · /u/koibKop4 ·

    DeepSeek V4 Flash 0731 appreciation post

    <!-- SC_OFF --><div class="md"><p>I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real.</p> <p>Everyday tasks with Hermes agent? Effortless.</p> <p>Coding tasks with OpenCode? I’m genuinely amazed at what it can handle. …

  38. r/LocalLLaMA TIER_1 English(EN) · /u/kuhunaxeyive ·

    Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks?

    <!-- SC_OFF --><div class="md"><p><em>(I am not a native speaker, written by myself, so please bear with me)</em></p> <p>I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence benchmarks. And those flaws render …

  39. r/LocalLLaMA TIER_1 Nederlands(NL) · /u/johnnyApplePRNG ·

    DeepSeek V4 Flash 0731 - ARC-AGI Results

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vi9zls/deepseek_v4_flash_0731_arcagi_results/"> <img alt="DeepSeek V4 Flash 0731 - ARC-AGI Results" src="https://external-preview.redd.it/sWhpb1GjRlbd3knWV_xC1C2WMQX5RRFImjvBgSF_7ZI.png?width=640&amp;crop=sma…

  40. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek-V4-Flash-0731, an MIT-licensed open-weight model, pulled 617.9k downloads in 30 days. That's real traction for a model that's not locked behind any API

    DeepSeek-V4-Flash-0731, an MIT-licensed open-weight model, pulled 617.9k downloads in 30 days. That's real traction for a model that's not locked behind any API. Click to see why it's climbing. https:// olud.ai/#leaderboard # OpenSource # AI # HuggingFace

  41. Mastodon — fosstodon.org TIER_1 Deutsch(DE) · [email protected] ·

    RT @UnslothAI: DeepSeek-V4-Flash can now be run locally 2× faster with DSpark! ⚡️ DSpark enables V4-Flash-0731 GGUFs to run ~1.4–2× faster

    RT @UnslothAI: DeepSeek-V4-Flash kann jetzt mit DSpark 2× schneller lokal ausgeführt werden! ⚡️ DSpark ermöglicht es V4-Flash-0731 GGUFs, ~1,4–2× schneller zu generieren, ohne Genauigkeitsverlust. DeepSeek-V4-Flash-0731 erreicht bis zu 120 Token/s. GGUFs: https:// huggingface.co/…

  42. r/LocalLLaMA TIER_1 English(EN) · /u/Ok_Ninja7526 ·

    Final optimization: from ~10 tok/s to ~15 tok/s on DeepSeek-V4-Flash-0731 at 128K ctx - 1 RTX 3090

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vh1qn3/final_optimization_from_10_toks_to_15_toks_on/"> <img alt="Final optimization: from ~10 tok/s to ~15 tok/s on DeepSeek-V4-Flash-0731 at 128K ctx - 1 RTX 3090" src="https://preview.redd.it/tj77v44nqqhh1…

  43. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📦 DeepSeek V4 Flash 0731 lands on the leaderboard with open weights: 1M context at $0.09 in / $0.18 out. A new option for long-context tasks without the high co

    📦 DeepSeek V4 Flash 0731 lands on the leaderboard with open weights: 1M context at $0.09 in / $0.18 out. A new option for long-context tasks without the high cost. https:// olud.ai/latest.html # AI # LLM # OpenSource

  44. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek V4 Flash is no longer text-only 👀 In our internal benchmarks, it delivered significantly better price-performance than others in its class. We’ve added

    DeepSeek V4 Flash is no longer text-only 👀 In our internal benchmarks, it delivered significantly better price-performance than others in its class. We’ve added vision for screen-level understanding, capabilities needed for WebBrain Available on HF: https:// huggingface.co/webbra…

  45. r/LocalLLaMA TIER_1 English(EN) · /u/dangerous_inference ·

    DeepSeek-V4-Flash on SM89 4x48gb 4090s with DSpark

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vfrjwl/deepseekv4flash_on_sm89_4x48gb_4090s_with_dspark/"> <img alt="DeepSeek-V4-Flash on SM89 4x48gb 4090s with DSpark" src="https://external-preview.redd.it/ODFwYWZ6b2M2Z2hoMVEbeTvpmi808aYcwKkY2JtviqN_2flas…

  46. r/LocalLLaMA TIER_1 Svenska(SV) · /u/giveen ·

    DeepSeek-v4-Flash-Mini 54GB GGUF running at ~20.5 t/s

    <!-- SC_OFF --><div class="md"><p>Took the REAP adaptation of DeepSeek-V4-Flash (<code>0xSero/DeepSeek-V4-Flash-0731-REAP</code>) along with <code>antirez/deepseek-v4-gguf</code> as inspiration, and decided to see how aggressive we could get with standard quant tricks to create a…

  47. r/LocalLLaMA TIER_1 Deutsch(DE) · /u/returnity ·

    DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vfhqkm/deepseek_v4_flash_vs_qwen3627b_35122b_and_gemma_4/"> <img alt="DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark" src="https://preview.redd.it/xucnql0lcehh1.png?width=140&amp;heigh…

  48. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🤖 DeepSeek's new V4 Flash really is 99% cheaper than Claude Opus 4.8: $0.28 vs $25 per million output tokens. It also beat Claude on Arena's front-end coding le

    🤖 DeepSeek's new V4 Flash really is 99% cheaper than Claude Opus 4.8: $0.28 vs $25 per million output tokens. It also beat Claude on Arena's front-end coding leaderboard and scores 82.7 on Terminal-Bench, ahead of some Claude tiers. That part is true. The nuance: on overall intel…

  49. r/LocalLLaMA TIER_1 Nederlands(NL) · /u/Gohab2001 ·

    Deepseek V4 flash 0731 ranks #21 on Agent Arena

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vff300/deepseek_v4_flash_0731_ranks_21_on_agent_arena/"> <img alt="Deepseek V4 flash 0731 ranks #21 on Agent Arena" src="https://preview.redd.it/522fsdwvtdhh1.png?width=140&amp;height=140&amp;auto=webp&amp;s=…

  50. r/LocalLLaMA TIER_1 English(EN) · /u/mrstoatey ·

    DeepSeek V4 Flash 0731 (Q4) now reaches 1,328 tok/s prefill and ~29 tok/s decode on one RTX PRO 6000

    <!-- SC_OFF --><div class="md"><p>I've been working on speeding up DeepSeek-V4-Flash-0731 in Krasis and have now got the long-prompt prefill quite a bit faster on a single RTX PRO 6000 96GB.</p> <p>These are timing-disabled internal Krasis results using INT4 experts. They aren't …

  51. r/LocalLLaMA TIER_1 English(EN) · /u/grumd ·

    Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

    <!-- SC_OFF --><div class="md"><p>I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the benchmark because it's quick to run, is pretty …

  52. r/LocalLLaMA TIER_1 English(EN) · /u/BlackBeardAI ·

    [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

    <!-- SC_OFF --><div class="md"><p>First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that:</p> <p><a href="https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_284b_moe_at_33_toks_single_68/"…

  53. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek-V4-Flash-0731 is pulling 236.1k monthly downloads—real demand, not just hype. That’s 2.1k stars on an MIT-licensed text model. The open-weight leaderbo

    DeepSeek-V4-Flash-0731 is pulling 236.1k monthly downloads—real demand, not just hype. That’s 2.1k stars on an MIT-licensed text model. The open-weight leaderboard shifts daily, and this one’s climbing for a reason. https:// olud.ai/#leaderboard # OpenSource # AI # HuggingFace

  54. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek V4 Flash on a Single AMD MI300X https:// github.com/ryanzhou/deepseek-v 4-flash-mi300x Comments: https:// news.ycombinator.com/item?id=4 9166386 # Hack

    DeepSeek V4 Flash on a Single AMD MI300X https:// github.com/ryanzhou/deepseek-v 4-flash-mi300x Comments: https:// news.ycombinator.com/item?id=4 9166386 # HackerNews # DeepSeek # V4 # Flash # AMD # MI300X # AI # Technology # MachineLearning # GPU

  55. r/LocalLLaMA TIER_1 English(EN) · /u/AbbreviationsSad5582 ·

    DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_284b_moe_at_33_toks_single_68/"> <img alt="DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config" src="https:…

  56. r/LocalLLaMA TIER_1 English(EN) · /u/LimpComedian1317 ·

    We tested Deepseek v4 flash, GLM 5.2, and Kimi K3 on hard agentic tasks, and DeepSeek just crushed

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1venecp/we_tested_deepseek_v4_flash_glm_52_and_kimi_k3_on/"> <img alt="We tested Deepseek v4 flash, GLM 5.2, and Kimi K3 on hard agentic tasks, and DeepSeek just crushed" src="https://preview.redd.it/uybjxyypj…

  57. r/LocalLLaMA TIER_1 English(EN) · /u/mintybadgerme ·

    I CANNOT believe I've got DeepSeek-V4-Flash-0731, a frontier model, running on my home PC. Insane!

    <!-- SC_OFF --><div class="md"><p>So this is the stuff of absolute insanity. In less than 20 months we've gone from super expensive cloud models only, to being able to run a Q3 quant of DeepSeek on an Intel Windows PC with a very average 24GB of VRAM. No wonder the big boys are p…

  58. r/LocalLLaMA TIER_1 English(EN) · /u/reto-wyss ·

    DeepSeek V4 Flash 0731 - Happy Numbers (700pp/18tg) and Thoughts

    <!-- SC_OFF --><div class="md"><p>Originally, I was only getting around 140pp/s and about 21tg/s, but the config with <code>-b 8192 -ub 8192 --cpu-moe</code> is vastly superior, let's say <strong>700pp/s</strong> and <strong>18tg/s</strong> in the most relevant range.</p> <p><str…

  59. Mastodon — fosstodon.org TIER_1 Deutsch(DE) · [email protected] ·

    RT @runsonai: I've been using DeepSeek V4 Flash all day and am very impressed. As a Hermes agent, it does everything you could wish for: agentic

    RT @runsonai: Ich habe den ganzen Tag über DeepSeek V4 Flash genutzt und bin sehr beeindruckt. Als Hermes-Agent erledigt er alles, was man sich wünscht: agentices Verhalten, Tool-Calling, Programmierung und Reaktionsfähigkeit. Die Schwäche von DS4F liegt im Harness; das Modell se…

  60. dev.to — LLM tag TIER_1 English(EN) · AI OpenFree ·

    Inside DeepSeek-V4-Flash-0731: 72,317 tensors of a harness-optimized MoE

    <p>On 2026-07-31, DeepSeek quietly shipped <strong>V4-Flash-0731</strong> under an MIT license. Instead of reading the model card, VIDRAFT's <strong>Darwin</strong> model-inspection platform read the actual <code>config.json</code> and every weight shard — <strong>72,317 tensors …

  61. r/LocalLLaMA TIER_1 English(EN) · /u/coder543 ·

    DeepSeek-V4-Flash-0731: When Low is higher than High

    <!-- SC_OFF --><div class="md"><p>I decided to test a few questions against DeepSeek-V4-Flash-0731. Locally, I was running Unsloth's UD-Q2_K_XL quant. After I saw the surprising shape of the results, I tested against DeepSeek's official API to confirm that I didn't do anything wr…

  62. r/LocalLLaMA TIER_1 English(EN) · /u/fragment_me ·

    Deepseek v4 flash - 100-150 faster t/s in prefill/pp.

    <!-- SC_OFF --><div class="md"><p>You have two choices here (in order of pref):</p> <ol> <li><p>Downgrade CUDA from 13.3 to 13.1 (skip 13.2 due to bugs) &lt;- prefer this (thanks to <a href="/u/fairydreaming">u/fairydreaming</a> for pointing this out)</p></li> <li><p>Use this vib…

  63. r/LocalLLaMA TIER_1 English(EN) · /u/Puzzleheaded_Base302 ·

    Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization - AI's narrative

    <!-- SC_OFF --><div class="md"><p># Running DeepSeek-V4-Flash-0731 (155 GB MoE) on a DGX Spark with vLLM-Moet 2-bit quantization</p> <p>I used Deepseek-v4-Flash-0731 cloud API settig up vllm-moet to run deepseek-v4-flash with MTP locally on single DGX Spark at 2-bit quant. Though…

  64. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek V4 Flash (0731) is quite close to GPT-5.6 Luna in terms of intelligence-cost ratio. # ai # news

    DeepSeek V4 Flash (0731) is quite close to GPT-5.6 Luna in terms of intelligence-cost ratio. # ai # news

  65. r/LocalLLaMA TIER_1 English(EN) · /u/Blahblahblakha ·

    DeepSeek-V4-Flash 284B on 5.3GB of memory

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vdbix4/deepseekv4flash_284b_on_53gb_of_memory/"> <img alt="DeepSeek-V4-Flash 284B on 5.3GB of memory" src="https://external-preview.redd.it/dmlxOHY4ZDh5d2doMTKeaIgiLAILwUoBdScnKpKTXpOnyWsUOZKXl9qB_yy6.png?wid…

  66. dev.to — LLM tag TIER_1 English(EN) · Hunter G ·

    DeepSeek V4-Flash Update Delivers Top Scores at Ultra-Low Prices

    <p>DeepSeek shipped the official V4-Flash on July 31 — <strong>open weights, MIT license, and a technical report, all on day one.</strong></p> <p>The striking part isn't that the score is high. It's that <strong>the score and the price arrived together</strong>: 50 on the Artific…

  67. r/LocalLLaMA TIER_1 English(EN) · /u/Hyungsun ·

    DeepSeek-V4-Flash-0731 UD-IQ3_XXS about 11t/s on 1x 7900 XTX 24GB + 3x MI60 32GB + 128GB DDR4

    <!-- SC_OFF --><div class="md"><p>Hello, Also I want to join the hype of posting token specs.</p> <p>CPU: 2x Intel Xeon CPU E5-2650 v4 @ 2.20GHz</p> <p>RAM: 2x 4 Channel 2400MHz DDR4</p> <p>GPU: 1x AMD Radeon 7900 XTX 24GB</p> <p>3x AMD Instinct MI60 32GB</p> <p>Strange GPU combi…

  68. Mastodon — fosstodon.org TIER_1 Deutsch(DE) · [email protected] ·

    RT @deepseek_ai: 🚀 The official DeepSeek-V4 Flash API is now available in public beta! 🔷 We have massively improved their agent capabilities – the

    RT @deepseek_ai: 🚀 Die offizielle DeepSeek-V4-Flash-API ist jetzt in der öffentlichen Beta verfügbar! 🔷 Wir haben ihre Agent-Fähigkeiten massiv verbessert – die Benchmark-Werte übertreffen nun die V4-Pro-Preview bei weitem. Entdecken Sie den enormen Leistungssprung unten! 👇 🔷 Die…

  69. r/LocalLLaMA TIER_1 English(EN) · /u/USBhost ·

    DeepSeek-V4-Flash-0731 UD-Q8_K_XL 17.20~ t/s on A6000 + 256GB DDR4

    <!-- SC_OFF --><div class="md"><p>Hello everyone I want to join the hype of posting specs.</p> <p>CPU: AMD EPYC 74F3 24-Core</p> <p>RAM: 8 Channel 3200 DDR4</p> <p>GPU: RTX A6000 48GB</p> <p>Prompt processing is in the high 70t/s (got down to mid 30t/s at 300k context). Inference…

  70. r/LocalLLaMA TIER_1 English(EN) · /u/HockeyDadNinja ·

    Expert-only IQ3 requant of DeepSeek-V4-Flash-0731: better KLD than UD-IQ3_S, 1.4x decode on a CPU-spill rig

    <!-- SC_OFF --><div class="md"><p>Hey all,</p> <p>tldr / who this helps: you run a mixed multi-GPU box where the experts spill to RAM, and you want to stay in the 3-bit tier instead of dropping to Q2 to make it fit. </p> <p><a href="https://huggingface.co/TacoTakumi/DeepSeek-V4-F…

  71. r/LocalLLaMA TIER_1 English(EN) · /u/Ok_Ninja7526 ·

    DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vcz61x/deepseekv4flash0731_udiq3_s_125_toks_on_rtx_3090/"> <img alt="DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5" src="https://external-preview.redd.it/NHV0ZnJ5NHB5dGdoMVngZSdHf_oCzXnIj…

  72. r/LocalLLaMA TIER_1 English(EN) · /u/esw123 ·

    DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vcrd6d/deepseek_v4_flash_0731_iq2_m_benchmark_for_dual/"> <img alt="DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s." src="https://preview.redd.it/3mzcq9labsgh1.png?width=640&amp…

  73. dev.to — LLM tag TIER_1 English(EN) · OctoLab ·

    How I Use DeepSeek V4 Flash: Reserve the Strongest Model for Uncertainty

    <p>I believe the group-chat comparison. I do not believe it reflects the model’s true value.</p> <p>One person used the same prompt with GLM 5.2 + Claude Code and DeepSeek V4 Flash + Codex. The first run produced a playable game in a little over twenty turns. Another similar test…

  74. Mastodon — fosstodon.org TIER_1 Italiano(IT) · [email protected] ·

    🧠 DeepSeek released DeepSeek V4 Flash 0731 via official API in public beta, with an update focused primarily on agentic capabilities. 👉

    🧠 # DeepSeek ha rilasciato DeepSeek V4 Flash 0731 tramite API ufficiale in public beta, con un aggiornamento concentrato soprattutto sulle capacità agentiche. 👉 I dettagli: https://www. linkedin.com/posts/alessiopoma ro_deepseek-claude-opus-share-7489318228167397376-DKwe/ ___ ✉️ …

  75. dev.to — LLM tag TIER_1 English(EN) · lora ·

    How I Added Low-Token Vision to DeepSeek V4 Flash

    <p>Instead of replacing the main model with a more expensive multimodal model, I gave a text-only agent an on-demand visual sensor.</p> <p>Many coding agents can already read repositories, write code, execute commands, and run tests. But in real development workflows, they often …

  76. dev.to — LLM tag TIER_1 English(EN) · TokenPAPA ·

    DeepSeek V4 Flash Official Release (0731): Agent Benchmarks Jump Past V4 Pro — Live on TokenPAPA

    <h1> DeepSeek V4 Flash Official Release (0731): Agent Benchmarks Jump Past V4 Pro </h1> <p>On July 31, 2026, DeepSeek officially released <strong>DeepSeek-V4-Flash</strong> to public beta. The API calling convention is unchanged — set <code>model</code> to <code>deepseek-v4-flash…

  77. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek's V4-Flash API launch improves agent capabilities with a post-training update, +25.8 on Terminal-Bench, and open-weights release. Source: Latent Space

    DeepSeek's V4-Flash API launch improves agent capabilities with a post-training update, +25.8 on Terminal-Bench, and open-weights release. Source: Latent Space https://www. latent.space/p/ainews-not-much -happened-today-038 # AI # Automation

  78. Mastodon — fosstodon.org TIER_1 Italiano(IT) · [email protected] ·

    DeepSeek has released the public beta of # V4FlashAPI. The new model shows capabilities for agents and benchmarks superior to V4 Pro Preview, marking a

    DeepSeek ha rilasciato la beta pubblica della # V4FlashAPI . Il nuovo modello mostra capacità per agenti e benchmark superiori alla V4 Pro Preview, segnando un passo avanti per l'ecosistema # AI e lo sviluppo di agenti intelligenti.

  79. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek has upgraded its V4-Flash model with major gains in agentic and coding capabilities. The 284B-parameter MoE model now includes DSpark speculative decod

    DeepSeek has upgraded its V4-Flash model with major gains in agentic and coding capabilities. The 284B-parameter MoE model now includes DSpark speculative decoding and supports the Responses API format. Pricing starts at 0.14 USD per million input tokens. https://www. marktechpos…

  80. r/LocalLLaMA TIER_1 English(EN) · /u/challis88ocarina ·

    DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf

    <!-- SC_OFF --><div class="md"><p>Antirez stealthily uploaded the new weights in the old folder... and there we were tapping our fingers.</p> <p><a href="https://huggingface.co/antirez/deepseek-v4-gguf/tree/main">https://huggingface.co/antirez/deepseek-v4-gguf/tree/main</a></p> <…

  81. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek V4 Flash now defaults to updated weights on AI Gateway, delivering a 25.8-point Terminal-Bench boost to 82.7 for stronger agentic coding tasks. No mode

    DeepSeek V4 Flash now defaults to updated weights on AI Gateway, delivering a 25.8-point Terminal-Bench boost to 82.7 for stronger agentic coding tasks. No model ID or code changes needed. Zero Data Retention providers arrive next week. Source: Vercel Blog https:// vercel.com/cha…

  82. r/LocalLLaMA TIER_1 English(EN) · /u/davidthesong ·

    Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and >50x cheaper

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vc4041/deepseek_v4_flash_is_now_2_open_weight_model_to/"> <img alt="Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and &gt;50x cheaper" src="https://preview.redd.it/h7zv5tb3tmgh1.png?width=140&amp;…

  83. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    DeepSeek's V4 Flash model now scores 50 on the AI Index, nearly matching GPT-5.6 Luna at 60% lower cost. A clear win for efficiency. Source: The Decoder AI http

    DeepSeek's V4 Flash model now scores 50 on the AI Index, nearly matching GPT-5.6 Luna at 60% lower cost. A clear win for efficiency. Source: The Decoder AI https:// the-decoder.com/new-deepseek-f lash-model-matches-openais-gpt-5-6-luna-at-roughly-60-percent-lower-cost/ # AI # Aut…

  84. r/LocalLLaMA TIER_1 English(EN) · /u/curiousily_ ·

    Initial testing of DeepSeek v4 Flash shows significant improvements in UI/UX design capabilities (despite being token hungry)

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vbxoi5/initial_testing_of_deepseek_v4_flash_shows/"> <img alt="Initial testing of DeepSeek v4 Flash shows significant improvements in UI/UX design capabilities (despite being token hungry)" src="https://exter…

  85. r/LocalLLaMA TIER_1 English(EN) · /u/RunawayPeeko ·

    DeepSeek V4 Flash unsloth quants are out!

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vbxk53/deepseek_v4_flash_unsloth_quants_are_out/"> <img alt="DeepSeek V4 Flash unsloth quants are out!" src="https://external-preview.redd.it/wMZUsudfTHG534WgBRWwhtsK20jasRrxNGwenbWxNxM.png?width=640&amp;crop…

  86. dev.to — LLM tag TIER_1 Nederlands(NL) · David ·

    DeepSeek V4 Flash vs V4 Pro: A Developer Decision Guide

    <p>With today's V4 Flash 0731 release, DeepSeek's V4 family now has two very different members, and picking between them is a real decision. Short version: Flash for agents and code, Pro for the hardest reasoning, Flash by forfeit if you want to run it yourself.</p> <h2> The spec…

  87. r/LocalLLaMA TIER_1 Nederlands(NL) · /u/Nyghtbynger ·

    Some DeepSeek-V4-Flash 20260731 opinion review

    <!-- SC_OFF --><div class="md"><p>First of all, I want to apologize if it's off-topic or in the wrong format.</p> <p>Having tried Deepseek Flash with reasoning high on a conceptually difficult task, involving Machine Learning classifiers and graphs. I am extremely impressed. It's…

  88. r/LocalLLaMA TIER_1 English(EN) · /u/SnooBunnies8392 ·

    DeepSeek-V4-Flash-0731 now far surpassing the DeepSeek-V4-Pro-Preview in benchmarks

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vbkvau/deepseekv4flash0731_now_far_surpassing_the/"> <img alt="DeepSeek-V4-Flash-0731 now far surpassing the DeepSeek-V4-Pro-Preview in benchmarks" src="https://preview.redd.it/bq9d2c2vyigh1.jpeg?width=640&am…

  89. r/LocalLLaMA TIER_1 English(EN) · /u/MagicZhang ·

    New DeepSeek V4-Flash achieves 50 on ArtificalAnalysis Index, 1 point below GLM-5.2 and GPT-5.6 Luna

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vbk5ob/new_deepseek_v4flash_achieves_50_on/"> <img alt="New DeepSeek V4-Flash achieves 50 on ArtificalAnalysis Index, 1 point below GLM-5.2 and GPT-5.6 Luna" src="https://preview.redd.it/mtmrp4lnrigh1.jpeg?wi…

  90. r/LocalLLaMA TIER_1 English(EN) · /u/Nunki08 ·

    DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vbidkp/deepseekv4flash_has_been_updated_the_official/"> <img alt="DeepSeek-V4-Flash has been updated, &quot;The official release of DeepSeek-V4-Pro will follow soon&quot;" src="https://preview.redd.it/mbz7sdw…

  91. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @jun_song: TRANSLATION: SuperDeepseek-V4-Flash runs on 2xDGX Spark with a peak of 122 tokens per second. The work is based on the recipe of

    RT @jun_song: TRANSLASATION: SuperDeepseek-V4-Flash läuft auf 2xDGX Spark mit einem Spitzenwert von 122 Token pro Sekunde. Die Arbeit basiert auf dem Rezept von @MiaAIlab zur Ausführung mit DFlash MTP. 4-Bit-Quantisierung mit weniger als ~1 % Qualitätsverlust. Außerdem wurde es a…

  92. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @pupposandro: Lucebox Engine now runs DeepSeek V4 Flash 0731 from a single 98.29 GB GGUF model on a 128 GB AMD Strix Halo System. The

    RT @pupposandro: Lucebox Engine betreibt nun DeepSeek V4 Flash 0731 von einem einzelnen 98,29 GB großen GGUF-Modell auf einem 128 GB AMD Strix Halo System. Das Modell erzielt 82/92 Punkte auf ds4-eval-92 und erreicht 32,7 Token pro Sekunde mit DSpark. Über unsere festgelegten Eva…

  93. Mastodon — mastodon.social TIER_1 English(EN) · sipirtu ·

    DeepSeek’s V4 Flash model lists at $0.14 per million input tokens and $0.28 per million output tokens, yet third parties undercut this. A 30x price hike may sti

    DeepSeek’s V4 Flash model lists at $0.14 per million input tokens and $0.28 per million output tokens, yet third parties undercut this. A 30x price hike may still leave it the cheapest for users with stable prompts due to its cache architecture. Source: Pandaily https:// pandaily…

  94. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @MiaAI_lab: Give me a better setup that runs DeepSeek v4 Flash locally at 80+ tokens/s. All you need is 2 DGX Sparks. I'm waiting for that

    RT @MiaAI_lab: Gebt mir ein besseres Setup, das DeepSeek v4 Flash lokal mit 80+ Token/s ausführt. Alles was ihr braucht sind 2 DGX Sparks. Ich warte darauf, dass Mike (@mike64t) Nvidia lacht, während die Leute auf einen 4.000 USD teuren glorifizierten Raspberry Pi hereinfallen, d…

  95. Mastodon — mastodon.social TIER_1 English(EN) · aisyndicate ·

    284B in 2 Bit auf 128 GB: DeepSeek V4 Flash mit 20 tok/s DeepSeek V4 Flash als 2-Bit-GGUF auf 128 GB Unified Memory: 20 tok/s, 81/100 im Tool-Eval. Ein validier

    284B in 2 Bit auf 128 GB: DeepSeek V4 Flash mit 20 tok/s DeepSeek V4 Flash als 2-Bit-GGUF auf 128 GB Unified Memory: 20 tok/s, 81/100 im Tool-Eval. Ein validiertes Laborobjekt für Offline-Betrieb, nicht für Effizienz. https:// aisyndicate.ch/deepseek-v4-fla sh-dgx-spark-test # AI…

  96. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    DeepSeek V4 Flash 0731’s price fell 36% in 30 days: $0.28→$0.18 per 1M output tokens. Open weights make it an interesting alternative now. https:// olud.ai/mode

    DeepSeek V4 Flash 0731’s price fell 36% in 30 days: $0.28→$0.18 per 1M output tokens. Open weights make it an interesting alternative now. https:// olud.ai/model/deepseek-v4-flas h-0731.html # AI # LLM # Pricing

  97. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @iam_elias1: Deepseek v4 Flash is overloaded because too many people are using it. This happens when you make a model too good. OpenCode (@opencode) has the

    RT @iam_elias1: Deepseek v4 Flash ist überlastet, weil zu viele Menschen es nutzen. Das passiert, wenn man ein Modell zu gut macht. OpenCode (@opencode) hat derzeit Kapazitätsprobleme bei Deepseek Flash aufgrund des beispiellosen Volumens; Sie können Fehler sehen - wir arbeiten a…

  98. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    I am just trying out DeepSeek v4 Flash 0731 and holly shit this thing is incredible for how small this model is and how little resources it needs. It only needs

    I am just trying out DeepSeek v4 Flash 0731 and holly shit this thing is incredible for how small this model is and how little resources it needs. It only needs 128 GB... which is at least 4 to 15 times less then the comparable new Models like GLM-5.2 or Kimi-K3... I guess China …

  99. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    DeepSeek V4 Flash on a Single AMD MI300X Article URL: https:// github.com/ryanzhou/deepseek-v 4-flash-mi300x Comments URL: https:// news.ycombinator.com/item?id

    DeepSeek V4 Flash on a Single AMD MI300X Article URL: https:// github.com/ryanzhou/deepseek-v 4-flash-mi300x Comments URL: https:// news.ycombinator.com/item?id=4 9166386 Points: 12 # Comments: 0 https:// github.com/ryanzhou/deepseek-v 4-flash-mi300x # Tech # Technology # TechNew…

  100. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @runsonai: I've been using DeepSeek V4 Flash all day and am very impressed. As a Hermes agent, it does everything you could wish for: agentic

    RT @runsonai: Ich habe den ganzen Tag über DeepSeek V4 Flash genutzt und bin sehr beeindruckt. Als Hermes-Agent erledigt er alles, was man sich wünscht: agentices Verhalten, Tool-Calling, Programmierung und Reaktionsfähigkeit. Die Schwäche von DS4F liegt im Harness; das Modell se…

  101. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @Bassmaster187: I have now tried DeepSeek V4 Flash 0731, as I have reached the limits with Qwen 3.6 27B. The new DeepSeek 0731 is supposed to be according to Benc

    RT @Bassmaster187: Ich habe jetzt DeepSeek V4 Flash 0731 ausprobiert, da ich mit Qwen 3.6 27B an die Grenzen gestoßen bin. Das neue DeepSeek 0731 soll laut Benchmarks extrem gut sein und extrem günstig. Ich hatte mal ein Angular-Plugin für Grafana geschrieben, und Angular ist jet…

  102. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @MiaAI_lab: DeepSeek v4 Flash feels like the Qwen3.6 27B moment all over again, just on a much larger scale. Never before has so much been possible for so little money

    RT @MiaAI_lab: DeepSeek v4 Flash fühlt sich wie der Qwen3.6 27B-Moment noch einmal an, nur auf viel größerer Ebene. Noch nie konnte man so viel für so wenig Geld tun. Dies könnte die Wirtschaftlichkeit von KI verändern. OpenCode (@opencode) DeepSeek Flash hat am 1. August 8 Billi…

  103. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @__tinygrad__: 245 Single-User-Tokens per second on DeepSeek-V4-Flash-0731, using only 2 of the RTX 6000 Blackwell GPUs! In honor of DeepSe

    RT @__tinygrad__: 245 Single-User-Tokens pro Sekunde auf DeepSeek-V4-Flash-0731, und dabei werden nur 2 der RTX 6000 Blackwell-GPUs genutzt! Zu Ehren von DeepSeek starten wir die 2-GPU-Edition unserer tinybox. Alle Hardwarekomponenten für die Installation von 4 GPUs sind enthalte…

  104. r/Anthropic TIER_1 English(EN) · /u/hibzy7 ·

    DeepSeek V4 Flash API is 18x cheaper on input, 28x cheaper on output, and matches Opus 4.8. Time for Claude to atleast reduce sonnet pricing

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1vcbac2/deepseek_v4_flash_api_is_18x_cheaper_on_input_28x/"> <img alt="DeepSeek V4 Flash API is 18x cheaper on input, 28x cheaper on output, and matches Opus 4.8. Time for Claude to atleast reduce sonnet pricin…

  105. r/Anthropic TIER_1 English(EN) · /u/Icy-Investment407 ·

    Newly released Deepseek V4 Flash Official scores close to Claude Opus 4.8. Price: 0.18$ / 1mil OUTPUT tokens

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1vc9zj1/newly_released_deepseek_v4_flash_official_scores/"> <img alt="Newly released Deepseek V4 Flash Official scores close to Claude Opus 4.8. Price: 0.18$ / 1mil OUTPUT tokens" src="https://preview.redd.it/u…

  106. Mastodon — mastodon.social TIER_1 Italiano(IT) · AI_BEAR_NEWS ·

    🤖 DeepSeek Launches Public Beta API for V4 Flash China's Top Model Now Available with Enhanced Agentic Capabilities and Benchmarks Outperforming V4-Pro-Pr

    🤖 DeepSeek lancia API beta pubblica per V4 Flash Il modello di punta cinese ora disponibile con capacità agentiche potenziate e benchmark che superano V4-Pro-Preview. Aperta la strada per sviluppatori e aziende che cercano alternative economicamente accessibili. Fonte: Bloomberg …

  107. r/OpenAI TIER_2 English(EN) · /u/sirMoped ·

    New post train of DeepSeek v4 flash is out

    <table> <tr><td> <a href="https://www.reddit.com/r/OpenAI/comments/1vc527y/new_post_train_of_deepseek_v4_flash_is_out/"> <img alt="New post train of DeepSeek v4 flash is out" src="https://preview.redd.it/vsao0zgg3ngh1.png?width=640&amp;crop=smart&amp;auto=webp&amp;s=017d0be1788b6…

  108. r/singularity TIER_2 English(EN) · /u/DeArgonaut ·

    DeepSeek V4 Flash 0731 ARC-AGI-1 and 2

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1vj65p3/deepseek_v4_flash_0731_arcagi1_and_2/"> <img alt="DeepSeek V4 Flash 0731 ARC-AGI-1 and 2" src="https://preview.redd.it/o3yky8nrm7ih1.png?width=140&amp;height=65&amp;auto=webp&amp;s=c276ccfedaa1f9730ea…

  109. r/singularity TIER_2 English(EN) · /u/Hot_Example_4456 ·

    Weights of Deepseek v4 flash 0731 have been released!!!

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1vbphpz/weights_of_deepseek_v4_flash_0731_have_been/"> <img alt="Weights of Deepseek v4 flash 0731 have been released!!!" src="https://external-preview.redd.it/4wuKMSgR8Kkk0pgA3-HhBYFoQX79t2M-87LJSpJS8lU.png?…

  110. r/singularity TIER_2 English(EN) · /u/Boring_Aioli7916 ·

    DeepSeek-V4-Flash Official API is now LIVE in public beta! Massive upgrades for flash model.

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1vbm5m8/deepseekv4flash_official_api_is_now_live_in/"> <img alt="DeepSeek-V4-Flash Official API is now LIVE in public beta! Massive upgrades for flash model." src="https://preview.redd.it/z35fca23cjgh1.jpeg?w…