DevTools Logo

LLM Benchmarks

LLM Benchmarks

Compare LLM coding, reasoning, speed and cost benchmarks side-by-side and get a task-specific model recommendation

Task

Weighted toward LiveCodeBench, SWE-bench Lite and Aider — for agentic coding and refactoring work.

Filters

Usage volume (ROI)

Priority

Max qualityMax savings

Model Ranking (14 models)

Est. Monthly Cost is a total spend estimate at 10,000 requests/mo × 2,000 in + 500 out tokens — not a per-1M-token rate. Per-token list prices are in the In / 1M and Out / 1M columns.

Task rank#SelectModelLicenseBadges
1
DeepSeek V4 ProDeepSeek
8082
4478 tok/s
1.0M100
$0.43$0.87$13.05DeepSeek
2
DeepSeek V4 FlashDeepSeek
7560
82168 tok/s
1.0M100
$0.14$0.28$4.20DeepSeek
3
Qwen 3.7 MaxQwen
7479
4478 tok/s
1M100
$1.48$4.42$51.63Qwen-Research
4
GPT-OSS 120BOpenAI
7459
77148 tok/s
131K100
$0.037$0.17$1.59Apache-2.0
5
Claude Haiku 4.5Anthropic
7270
59108 tok/s
200K100
$1.00$5.00$45.00Proprietary
6
Gemma 4 31BGoogle
7256
77148 tok/s
262K100
$0.10$0.34$3.70Gemma
7
MiMo-V2.5Xiaomi
7256
77148 tok/s
1.1M100
$0.14$0.28$4.20Open-Weights-Other
8
Qwen 3.7 PlusQwen
7156
77148 tok/s
1M100
$0.32$1.28$12.80Qwen-Research
9
Claude Sonnet 5Anthropic
7080
4478 tok/s
1M100
$2.00$10.00$90.00Proprietary
10
Qwen 3.7 FlashQwen
6841
100210 tok/s
1M100
$0.030$0.13$1.25Qwen-Research
11
Nemotron 3 Ultra (free)NVIDIA
6741
96195 tok/s
1M100
FreeFree$0.00Llama-Community
12
Claude Opus 5Anthropic
5886
3360 tok/s
1M100
$5.00$25.00$225.00Proprietary
13
DevTools Code Quality (1.5B)DevTools (self-hosted, Qwen2.5-Coder base)
4332
028 tok/s
4K6
FreeFree$0.00Apache-2.0
14
DevTools Code Light (0.5B)DevTools (self-hosted, Qwen2.5-Coder base)
4324
2855 tok/s
4K6
FreeFree$0.00Apache-2.0

Pricing, context windows and model names come from the same catalog as the LLM token counter. Benchmark scores are hand-reviewed reference values based on public leaderboards (LM Arena, SWE-bench), not a live leaderboard feed — last updated 2026-08-20.

Examples

Find the best model for agentic coding

Open the Coding tab and read the Best Value badge — the ranking leans on LiveCodeBench, SWE-bench Lite and Aider scores at your chosen request volume.

Budget a high-volume chat feature

Switch to the Fast Chat tab, enter your expected monthly requests and average tokens per request, and slide the priority toward Savings.

Try a fully offline browser model

Open the Edge/Browser tab, check a WebLLM model's download size and VRAM requirement, then use its "Try in browser" link.

About this tool

Picking a model isn't just about a single leaderboard number — the best choice for an agentic coding task is rarely the best choice for a cheap high-volume chatbot or a fully offline browser tool. This matrix scores every model against the benchmarks that actually matter for five task types (coding, documentation/RAG, reasoning, edge/browser and fast chat), weighted differently per task, and rolls them into one comparable Task Score.

Pick a task tab to filter the table to relevant models and see the weighting behind the score. Enter your own monthly request volume and average input/output tokens to see an estimated monthly cost per model, and slide the quality-vs-savings priority to re-rank instantly. Select up to 4 models to compare on a radar chart across quality, cost, speed, context and (for edge models) memory footprint.

Every model card is source-linked (LM Arena, SWE-bench) and dated. Benchmark scores are hand-reviewed reference values, not a live leaderboard feed — this stays a static, client-side snapshot with no network calls in this phase.

How to use

  1. Pick a task

    Choose Coding, Documentation & RAG, Reasoning, Edge/Browser or Fast Chat — the weighting behind the score changes per task.

  2. Enter your traffic

    Type your expected monthly requests and average input/output tokens per request to see an estimated monthly cost per model.

  3. Slide the priority

    Move the quality-vs-savings slider to instantly re-rank models toward maximum quality or maximum cost savings.

  4. Compare on a radar chart

    Check up to 4 models and open Compare to see quality, cost, speed, context and memory plotted side-by-side.

Use cases

Model selection for a new feature

Compare quality, speed and cost together instead of chasing a single leaderboard number that doesn't match your task.

Budgeting a chatbot or agent

Enter real traffic assumptions to see estimated monthly cost per model before committing to a provider.

Evaluating on-device/offline AI

Filter to Edge/Browser to compare WebLLM models on download size, VRAM and quality trade-offs.

Common mistakes

Mistake:Comparing raw benchmark numbers across different task types

Fix:Use the task-specific Task Score instead — it's normalized and weighted for what actually matters for that kind of work.

Mistake:Treating this as a live, auto-updating leaderboard

Fix:Scores are a dated, hand-reviewed snapshot (see the last-updated note) — re-check official sources for the latest numbers before a high-stakes decision.

Frequently asked questions

References & standards