DevTools Logo

LLM Benchmarks

LLM Benchmarks

Compare LLM coding, reasoning, speed and cost benchmarks side-by-side and get a task-specific model recommendation

Task

Weighted toward LiveCodeBench, SWE-bench Lite and Aider — for agentic coding and refactoring work.

Filters

Usage volume (ROI)

Priority

Max qualityMax savings

Model Ranking (14 models)

Est. Monthly Cost is a total spend estimate at 10,000 requests/mo × 2,000 in + 500 out tokens — not a per-1M-token rate. Per-token list prices are in the In / 1M and Out / 1M columns.

Task rank#SelectModelLicenseBadges
1
DeepSeek V4 ProDeepSeek
8082
4478 tok/s
1.0M100
$0.43$0.87$13.05DeepSeek
2
DeepSeek V4 FlashDeepSeek
7560
82168 tok/s
1.0M100
$0.14$0.28$4.20DeepSeek
3
Qwen 3.7 MaxQwen
7479
4478 tok/s
1M100
$1.48$4.42$51.63Qwen-Research
4
GPT-OSS 120BOpenAI
7459
77148 tok/s
131K100
$0.037$0.17$1.59Apache-2.0
5
Claude Haiku 4.5Anthropic
7270
59108 tok/s
200K100
$1.00$5.00$45.00Proprietary
6
Gemma 4 31BGoogle
7256
77148 tok/s
262K100
$0.10$0.34$3.70Gemma
7
MiMo-V2.5Xiaomi
7256
77148 tok/s
1.1M100
$0.14$0.28$4.20Open-Weights-Other
8
Qwen 3.7 PlusQwen
7156
77148 tok/s
1M100
$0.32$1.28$12.80Qwen-Research
9
Claude Sonnet 5Anthropic
7080
4478 tok/s
1M100
$2.00$10.00$90.00Proprietary
10
Qwen 3.7 FlashQwen
6841
100210 tok/s
1M100
$0.030$0.13$1.25Qwen-Research
11
Nemotron 3 Ultra (free)NVIDIA
6741
96195 tok/s
1M100
FreeFree$0.00Llama-Community
12
Claude Opus 5Anthropic
5886
3360 tok/s
1M100
$5.00$25.00$225.00Proprietary
13
DevTools Code Quality (1.5B)DevTools (self-hosted, Qwen2.5-Coder base)
4332
028 tok/s
4K6
FreeFree$0.00Apache-2.0
14
DevTools Code Light (0.5B)DevTools (self-hosted, Qwen2.5-Coder base)
4324
2855 tok/s
4K6
FreeFree$0.00Apache-2.0

Pricing, context windows and model names come from the same catalog as the LLM token counter. Benchmark scores are hand-reviewed reference values based on public leaderboards (LM Arena, SWE-bench), not a live leaderboard feed — last updated 2026-08-20.