Frontier Model Benchmarks (Aug 2026)

Latest frontier flagship per lab · Chinese labs + NVIDIA · Excludes OpenAI, Anthropic, Google Gemini · MiMo & StepFun removed · + GLM-5.3-Flash, Qwen3.8-27B, Qwen3.8-Flash added. Numbers are vendor-reported from official READMEs/blogs unless marked as independent eval.

As of 4 Sep 2026 6 labs · 10 models GLM 5.3 + 5.3-Flash · Qwen 3.8-Max / 27B / Flash · Kimi K3 · DS 0813/0731 Sources: vendor · AA · Vals · Arena Agent · Next.js · LMArena · LiveBench Pricing: official API · USD / M tokens

Latest model per lab

MiniMax
M3
Jun 1 · 428B/23B act · 1M ctx · MSA · MiniMax license
$0.30 / $1.20 in · out
AA II 35.7 (Sep refresh) · weights released · Still frontier (no M4 yet)
DeepSeek
V4-Pro 0813
Aug 13 · 1.6T/49B act + DSpark · 1M ctx · MIT · supersedes Apr 24 Preview
$0.66 / $1.98 in · out (peak; $0.435/$0.87 off-peak cache-miss)
Official release · TB2.1 87.9 · DeepSWE 62.7 · HLE 42.7/60.0 w tools · HF 0813
DeepSeek
V4-Flash 0731
Jul 31 · 284B/13B act + DSpark · 1M ctx · MIT · supersedes Flash Preview
$0.35 / $0.88 in · out (est. flash tier)
Official Flash · TB2.1 82.7 · DeepSWE 54.4 · beats Pro Preview (72.1) · HF 0731
Moonshot
Kimi K3
Jul 17 · 2.8T/104B act (16/896) · 1M ctx · KDA + AttnRes · Kimi K3 License
$3.00 / $15.00 in · out
First 3T-class open · TB2.1 88.3 · GPQA 93.5 · BrowseComp 91.2 · HF K3
Z.ai
GLM-5.3
Aug 14 · 743B MoE base (same as 5.2) · PT upgrade · gated HF `zai-org/GLM-5.3`
Coding Plan ~$18/mo · API g.b. waitlist
Supersedes 5.2 · TB2.1 88.2 · TB3.0 28.3 · DeepSWE 66.9 · CyberGym 84.5 · openlm.ai
Z.ai
GLM-5.3-Flash
Aug 26 · 320B/18B act · hybrid sparse+linear attn · mHC · 1.31M ctx · native multimodal · MIT
$0.15 / $0.50 in · out
Open weights · TB2.1 84.3 · DeepSWE 63.4 · AutomationBench 48.8 · HLE w/tools 55.3 · AA II 57 · HF
Alibaba
Qwen3.8-Max
Aug 3 · 2.4T/95B act · 1M ctx · multimodal · open weights next week · API qwen3.8-max
$2.00 / $6.00 in · out
TB2.1 86.6 · SWE-Pro 67.7 · GPQA 92.6 · OSWorld-Ver. 86.1 · PaperBench 93.0 · qwen.ai
Alibaba
Qwen3.8-27B
Aug 14 · 27B dense (28B w/ vision) · Gated DeltaNet hybrid · 262K ctx (1M YaRN) · native VLM · Apache 2.0
open weights · 24GB GPU @ 4-bit
TB2.1 73.0 · SWE-Pro 61.7 · DeepSWE 42.2 · GPQA 89.2 · LCB v6 90.3 · OSWorld-Ver 84.3 · HF
Alibaba
Qwen3.8-Flash
Aug 26 · 125B/6B act MoE (512 experts) + 51B n-gram table · Qwen4 arch preview (QSA) · 256K→1M YaRN · Qwen Community Lic.
$0.15 / $0.47 in · out
Open preview = Flash-Next · SWE-Pro 62.5 · DeepSWE 58.7 · GPQA 91.7 · LCB v6 91.9 · NL2Repo 48.1 · HF
NVIDIA (US)
Nemotron 3 Ultra
Jun 4 · 550B/55B act · 1M ctx · OpenMDW · Mamba-Transf.
open weights · self-host / NIM
US open-weight leader, below Chinese frontier · AA II 29.3 (Sep refresh)

Removed: MiMo-V2.5-Pro & Step-3.7 Flash. Added Sep: GLM-5.3-Flash (Aug 26, MIT, 320B/18B), Qwen3.8-27B (Aug 14, Apache 2.0), Qwen3.8-Flash (Aug 26, Qwen4-arch preview). Previous K2.6 / GLM-5.2 / DS Preview / Qwen3.7 archived.

API pricing (USD per 1M tokens · official vendor)

Purple highlight = lowest in row Open-weight models also support self-hosting
Model Access Input (cache miss) Output Cached input Notes
MiniMax M3 API $0.30 $1.20 $0.06 ≤512k promo (list $0.60 / $2.40)
DeepSeek V4-Pro 0813 Open + API $0.66 $1.98 $0.0036 Peak; off-peak $0.435/$0.87 · MIT · +DSpark · docs
DeepSeek V4-Flash 0731 Open + API $0.35 $0.88 $0.0036 Flash tier · MIT · DSpark enabled · 1M ctx
Kimi K3 Open + API $3.00 $15.00 $0.60 Kimi K3 License · 1M ctx · platform.kimi.ai
GLM-5.3 Gated + API TBD TBD - Gated repo `zai-org/GLM-5.3` · Coding Plan · weights late Aug pending safety
GLM-5.3-Flash Open + API $0.15 $0.50 $0.03 320B/18B · MIT · 1.31M ctx · ~10x cheaper than 5.2 · z.ai
Qwen3.8-Max API + open $2.00 $6.00 $0.20 2.4T/95B MoE · 1M ctx · open weights next week · API qwen3.8-max
Qwen3.8-27B Open (Apache 2.0) - - - Self-host · 27B dense + vision · 262K ctx (1M YaRN) · 24GB GPU @ 4-bit · API "coming soon" on QwenCloud
Qwen3.8-Flash API + open preview $0.15 $0.47 - 125B/6B MoE + 51B n-gram · Qwen4 arch preview (Flash-Next open) · 256K→1M YaRN · QwenCloud
Nemotron 3 Ultra Open weights - - - OpenMDW-1.1 · self-host / NIM · 550B/55B MoE

Pricing cross-checked 2026-08-21 · DeepSeek 0813 peak pricing from AIToolsReview; off-peak $0.435 · K3 $3/$15 · Qwen3.8 $2/$6 official (vs Qwen3.7 $2.5/$7.5) · GLM-5.3 token pricing not yet on docs.z.ai/pricing

Independent composite scores

Artificial Analysis (independent) Vals.ai (independent)
Benchmark M3 DS Pro 0813 DS Flash 0731 K3 GLM-5.3 5.3-Flash Q3.8-Max Q3.8-27B Q3.8-Flash Nemotron
AA Intelligence Index ● 35.7 42.1 max 40.8 max 50.2 max 48.6 max 57 v4.1.1 46.9 52 xhigh - 29.3 AA CritPt / AA-LCR ● - - - 23.4 / 74.7 - - - - - - AA-Omniscience (hallucination) ● 1.35 0.83 - 19.7 14.3 7.47 4.32 - - -0.4 AA-Briefcase Elo ● 1096 1266 1260 1497 1515 1459 1392 - - - Vals Index ● 42.72% 52.37% 53.57% 57.81% 56.97% - 51.84% - - 27.39% GDPval-AA v2 Elo ● - 1496 - 1586 1678 1673 1630 - - - Arena Agent Net Improvement ● - - - - 4.37% ±2.48 (#10, 5.2) - - - - -

All AA rows pulled live from artificialanalysis.ai ld+json datasets (Sep 2026 index refresh — II rescaled vs Jul v4.1, same rank order). Vals Index from vals.ai Sep 4 2026 snapshot + vals.ai model page for DS Pro 0813 (52.37%, #18). Orange (AA) tags in the matrix below = Artificial Analysis independent runs filling vendor gaps.

Full benchmark matrix (vendor-reported unless noted · DeepSeek harness = minimal, max effort, temp 1.0 top_p 0.95 · K3 harness = Kimi Code, max)

Official model card / blog Green highlight = best in row among models DS internal: † DSBench
Benchmark M3 DS Pro 0813 DS Flash 0731 K3 GLM-5.3 5.3-Flash Q3.8-Max Q3.8-27B Q3.8-Flash Nemotron
Knowledge & Reasoning
GPQA Diamond 92.9 (AA) 92.8 (AA) 88.1 93.5 91.7 (AA) - 92.6 89.2 91.7 87.0
CritPt (AA) - 18.0 (AA) - 23.4 19.1 (AA) - 20.0 (AA) - - 3.1
AA-LCR 83.0 (AA) 80.3 (AA) - 88.7 (AA) - - 80.3 (AA) - - -
HLE (no tools) - 42.7 37.8 43.5 42.3 (AA) - 43.6 30.8 35.9 26.7
HLE w/ tools - 60.0 51.5 56.0 - 55.3 56.2 - - 37.4
MMLU-Pro - 87.5 86.2 - - - - - - 86.8
Coding & Agentic — Terminal / SWE
Terminal-Bench 2.1 66.0 (2.1) 87.9 82.7 88.3 88.2 84.3 86.6 73.0 - 56.4 (2.1)
Terminal-Bench 3.0 - - - - 28.3 - - - - -
SWE-bench Pro 59.0 55.4 52.6 - - - 67.7 61.7 62.5 -
SWE-bench Verified - 80.6 79.0 - - - - - - 70.7
NL2Repo - 61.5 54.2 - 58.0 56.3 55.9 42.3 48.1 -
DeepSWE (v1.1) - 62.7 54.4 67.5 66.9 63.4 56.6 42.2 58.7 -
ProgramBench - - - 77.8 19.0 - - - - -
FrontierSWE - - - 81.2 78.1 - 73.5 - - -
SWE-Marathon v1.1 - - - 42.0 42.5 - - - - -
PostTrainBench 0.37 (37%) - - 36.6 39.8 - - - - -
PaperBench - - - - - - 93.0 - - -
OSWorld-Verified - - - 84.8 - - 86.1 84.3 - -
IFBench 82.9 (AA) - - - - - 82.8 79.5 81.3 81.4 (AA)
OmniDocBench 1.5 - - - - - - 92.1 91.1 - -
MCPAtlas Public 74.2 73.6 - 84.2 - - - 31.1 - -
MCPMark Verified - - - 94.5 - - - - - -
CyberGym - 83.3 76.7 80.0 84.5 - - 43.2 - -
Toolathlon-Verified - 74.1 70.3 76.5 73.0 78.4 72.5 67.1 73.5 -
BrowseComp 83.5 - - 91.2 - - - - - 44.4
Agents' Last Exam - 25.7 25.2 28.3 28.5 26.3 27.0 20.4 24.3 -
AutomationBench Public - 31.8 25.1 30.8 - 48.8 27.3 - - -
DSBench-FullStack † - 71.1 68.7 73.7 - - - - - -
DSBench-Hard † - 67.2 59.6 63.0 - - - - - -
Agentic / Business & Other
GDPval-AA v2 (Elo) - 1496 (AA) - 1586 1678 - 1630 (AA) - - -
APEX-Agents - 38.3 33.0 41.0 - - - - - -
Kimi Code Bench 2.0 - - - 72.9 - - - - - -
MLS-Bench-Lite - - - 48.3 - - - - - -
SciCode - 51.0 (AA) - 58.7 59.0 (AA) - 54.1 (AA) - - 44.6
KernelBench Hard 28.8 - - - - - - - - -
Long Context
MRCR 1M / AA-LCR - 83.5 78.7 74.7 (AA-LCR) - - - - - -
CorpusQA 1M - 62.0 (Pro Preview) 60.5 - - - - - - -

DeepSeek 0813/0731 from HF 0813 & HF 0731 (minimal harness, max, temp 1.0/top_p 0.95). K3 from HF K3 (Kimi Code, max). GLM-5.3 from openlm.ai & Z.ai blog Aug 14. Qwen3.8-Max from qwen.ai/blog?id=qwen3.8 (Aug 3, 2.4T/95B). (AA) = Artificial Analysis independent run (artificialanalysis.ai/evaluations/*) where the vendor never published that benchmark. DSBench † = internal DeepSeek sets. Remaining blanks = no vendor AND no AA number exists.

Next.js agent evals (nextjs.org/evals · independent)

Vercel Next.js code gen & migration tasks Green = best success % · Orange = fastest time

Official harness measuring pass rate on Next.js generation/migration. Not comparable to SWE-bench vendor tables. AGENTS.md = extra passes with bundled docs.

Metric M3 DS Pro 0813 DS Flash 0731 K3 GLM-5.3 5.3-Flash Q3.8-Max Q3.8-27B Q3.8-Flash Nemotron
Success rate ● - - - - 75% (5.2, OpenCode) - - - - - Success w/ AGENTS.md 96% - - - - - - - - - Avg execution time 181.30s - - - - - - - - - Harness - - - - OpenCode (5.2) - - - - -

Only M3 row populated from Aug 2026 snapshot. K3 / 0813 / 0731 / 5.3 not yet on Next.js leaderboard — use bench matrix + Arena Agent for K3/5.3 agent signal.

Independent evals (third-party leaderboards)

LMArena Elo Arena Agent LiveBench OLLB EQ-Bench Green = best

LMArena via arena-ai-leaderboards 2026-06-02 snapshot. API-only Qwen columns keep LMArena. OLLB = open-weight only.

Metric M3 DS Pro 0813 DS Flash 0731 K3 GLM-5.3 5.3-Flash Q3.8-Max Q3.8-27B Q3.8-Flash Nemotron
LMArena Text Elo ● - - - - - - - - - - LMArena Code Elo ● - - - - - - - - - - LiveBench global avg ● - - - - - - - - - - Arena Agent Net Improvement ● - - - - 4.37% ±2.48 #10 (5.2) - - - - - Open LLM LB Average ● TBD - - - - - - - - - EQ-Bench Creative Writing Elo ● - - - - - - - - - -

No 0813/0731/K3/5.3 entries in LMArena top-20 or LiveBench Jun crawl yet. Check OLLB for HF `open-llm-leaderboard/contents` — none of the Aug weights appear yet (4576 entries checked Jun 17).

Sources