updated 24 Jul 2026

Open Source LLM Leaderboard

This open source LLM leaderboard displays the latest public benchmark performance for open-weight and open-source models released after April 2024. The data comes from model providers as well as independently run evaluations by Vellum or the open-source community. We feature results from non-saturated benchmarks, excluding outdated benchmarks (e.g. MMLU).

Top open source models per tasks

Best in Reasoning (GPQA Diamond)

95%92%88%85%81%
93.5
Kimi K3
93
MiniMax M3
91.2
GLM 5.2
90.5
Kimi K2.6
90.1
DeepSeek V4 Pro
Best in Reasoning (GPQA Diamond)
ModelScore
Kimi K393.5%
MiniMax M393%
GLM 5.291.2%
Kimi K2.690.5%
DeepSeek V4 Pro90.1%

Best in Agentic Coding (SWE Bench)

85%81%76%72%68%
80.6
DeepSeek V4 Pro
80.5
MiniMax M3
80.2
Kimi K2.6
79
DeepSeek V4 Flash
76.8
Kimi K2.5
Best in Agentic Coding (SWE Bench)
ModelScore
DeepSeek V4 Pro80.6%
MiniMax M380.5%
Kimi K2.680.2%
DeepSeek V4 Flash79%
Kimi K2.576.8%
New

Best in Computer Use (OSWorld)

75%72%69%66%63%
73.1
Kimi K2.6
70.1
MiniMax M3
Best in Computer Use (OSWorld)
ModelScore
Kimi K2.673.1%
MiniMax M370.1%
New

Best in Browsing (BrowseComp)

95%89%84%78%72%
91.2
Kimi K3
85.9
DeepSeek V4 Flash
83.5
MiniMax M3
83.4
DeepSeek V4 Pro
83.2
Kimi K2.6
Best in Browsing (BrowseComp)
ModelScore
Kimi K391.2%
DeepSeek V4 Flash85.9%
MiniMax M383.5%
DeepSeek V4 Pro83.4%
Kimi K2.683.2%
New

Best in Terminal Use (Terminal-Bench 2.1)

90%68%45%23%0%
88.3
Kimi K3
81
GLM 5.2
66
MiniMax M3
50.8
Kimi K2.5
35.7
Kimi K2 Thinking
Best in Terminal Use (Terminal-Bench 2.1)
ModelScore
Kimi K388.3%
GLM 5.281%
MiniMax M366%
Kimi K2.550.8%
Kimi K2 Thinking35.7%

Best in Visual Reasoning (ARC-AGI 2)

15%14%12%11%9%
12
Kimi K2.5
Best in Visual Reasoning (ARC-AGI 2)
ModelScore
Kimi K2.512%

Inference provider comparison

Fastest (Throughput)

2012150910065030
1828.8
Cerebras
749.2
Fireworks AI
592.9
Together AI
476.8
Groq
283
Baseten
Fastest (Throughput)
ProviderValue
Cerebras1828.8
Fireworks AI749.2
Together AI592.9
Groq476.8
Baseten283

Lowest Latency (TTFT)

43210
0.27s
Baseten
0.47s
Together AI
0.54s
Cerebras
0.7s
Groq
1.12s
Novita AI
3.43s
Fireworks AI
Lowest Latency (TTFT)
ProviderValue
Baseten0.27s
Together AI0.47s
Cerebras0.54s
Groq0.7s
Novita AI1.12s
Fireworks AI3.43s

Lowest Cost (per 1M tokens)InputOutput

$1$0.75$0.5$0.25$0
$0.05
$0.25
Novita AI
$0.1
$0.5
Baseten
$0.15
$0.6
Fireworks AI
$0.15
$0.6
Together AI
$0.15
$0.6
Groq
$0.35
$0.75
Cerebras
Lowest Cost (per 1M tokens)
ProviderInputOutput
Novita AI$0.05$0.25
Baseten$0.1$0.5
Fireworks AI$0.15$0.6
Together AI$0.15$0.6
Groq$0.15$0.6
Cerebras$0.35$0.75

Compare open source models

vs
Kimi K3GLM 5.2
Context size1,048,5761,000,000
Cutoff date-Mar 2026
I/O cost$3 / $15$0.95 / $3
Max output-128,000
Latency4.46s1.14s
Speed35.2 t/s347 t/s
Best Overall (HLE)
Kimi K3
56
GLM 5.2
54.7
Best in Terminal Use (Terminal-Bench 2.1)
Kimi K3
88.3
GLM 5.2
81
Best in Agentic Coding (SWE-Bench)
Kimi K3
76.8
GLM 5.2-
Best in Reasoning (GPQA Diamond)
Kimi K3
93.5
GLM 5.2
91.2

Model Comparison

ModelContext sizeCutoff dateI/O costMax outputLatencySpeed
Kimi K31,048,576-$3 / $15-4.46s35.2 t/s
GLM 5.21,000,000Mar 2026$0.95 / $3128,0001.14s347 t/s
Kimi K2.6256,000-$0.95 / $4-0.68s342.6 t/s
DeepSeek V4 Flash1000000Jan 2026$0.14 / $0.283840001.42s107.9 t/s
DeepSeek V4 Pro1000000Jan 2026$0.435 / $0.873840001.2s174.9 t/s
Kimi K2.5256,000Apr 2024$0.6 / $2.533,0000.69s337.7 t/s
DeepSeek-R1128,000Dec 2024$0.55 / $2.198,0001.18s30.1 t/s
Llama 3.1 405b128,000Dec 2023$3.5 / $3.540960.73s969 t/s
Llama 3.3 70b128,000July 2024$0.59 / $0.732,7680.52s2500 t/s
DeepSeek V3 0324128,000Dec 2024$0.27 / $1.18,0001.9s36.4 t/s
Qwen2.5-VL-32B131,000Dec 2024-8,000--
Gemma 3 27b128,000Nov 2024$0.07 / $0.0781920.72s59 t/s
Llama 4 Maverick10,000,000November 2024$0.2 / $0.68,0000.45s126 t/s
Llama 4 Scout10,000,000November 2024$0.11 / $0.348,0000.33s2600 t/s
MiniMax M31,048,576Mar 2026$0.6 / $2.4512,0000.85s98.6 t/s

Context window, cost and speed comparison

Models Context Window Input Cost / 1M tokens Output Cost / 1M tokens Speed (tokens/second) Latency
Kimi K31,048,576$3$1535.2 t/s4.46 seconds
GLM 5.21,000,000$0.95$3347 t/s1.14 seconds
Kimi K2.6256,000$0.95$4342.6 t/s0.68 seconds
DeepSeek V4 Flash1000000$0.14$0.28107.9 t/s1.42 seconds
DeepSeek V4 Pro1000000$0.435$0.87174.9 t/s1.2 seconds
Kimi K2.5256,000$0.6$2.5337.7 t/s0.69 seconds
DeepSeek-R1128,000$0.55$2.1930.1 t/s1.18 seconds
Llama 3.1 405b128,000$3.5$3.5969 t/s0.73 seconds
Llama 3.3 70b128,000$0.59$0.72500 t/s0.52 seconds
DeepSeek V3 0324128,000$0.27$1.136.4 t/s1.9 seconds
Qwen2.5-VL-32B131,000n/an/an/an/a
Gemma 3 27b128,000$0.07$0.0759 t/s0.72 seconds
Llama 4 Maverick10,000,000$0.2$0.6126 t/s0.45 seconds
Llama 4 Scout10,000,000$0.11$0.342600 t/s0.33 seconds
MiniMax M31,048,576$0.6$2.498.6 t/s0.85 seconds

Benchmark glossary

Humanity's Last Exam
A crowd-sourced exam of extremely hard questions spanning every academic discipline. Designed to be the final exam before superhuman AI.
GPQA Diamond
Graduate-level science questions curated by domain experts. Tests advanced reasoning across physics, chemistry, and biology.
SWE-Bench Verified
Real GitHub issues from popular Python repos that the model must resolve end-to-end. Measures agentic software engineering ability.
AutoBench
Automation benchmark evaluating a model's ability to complete real-world work automation tasks using tools and multi-step workflows.
OSWorld-Verified
Real-world computer use tasks requiring GUI interaction in desktop environments. Measures end-to-end task completion on a real OS.
BrowseComp
Agentic web search benchmark testing a model's ability to browse and extract information from the web to answer complex questions.
Terminal-Bench 2.1
Terminal and tool use benchmark evaluating a model's ability to execute multi-step tasks in a terminal environment.
ARC-AGI 2
Abstract visual puzzles requiring novel pattern recognition. Tests fluid intelligence and generalization beyond training data.

The Personal AI you were promised

GET STARTED