A recent benchmark of seven Chinese LLM APIs, based on over 1900 real-world API calls from May 2026, reveals that no single model excels across all tasks. DeepSeek-v4-pro achieved the highest overall score of 81.1, demonstrating strong performance in code generation and math reasoning, while Kimi K2.6 Thinking led in hallucination control with a score of 90.0. Doubao Seed2.0-pro topped code generation at 85.7, but models showed significant variation in token efficiency and response times, with some taking over 3 seconds to respond. The report also highlights the challenges of managing multiple API keys and suggests unified gateways like GoldBean can simplify access and potentially reduce costs. AI
IMPACT Highlights performance variations and cost efficiencies across Chinese LLM APIs, guiding developers in selecting optimal models for specific tasks.
RANK_REASON Benchmark report evaluating multiple LLM APIs.
- Bonree Data
- DeepSeek
- DeepSeek V4-Pro
- Doubao Seed2.0-pro
- General Language Model
- GoldBean
- Kimi K2.6 Thinking
- MIT
- OpenAI
- Qwen
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →