LLM Benchmarks
Compare LLM coding, reasoning, speed and cost benchmarks side-by-side and get a task-specific model recommendation
Task
Weighted toward LiveCodeBench, SWE-bench Lite and Aider — for agentic coding and refactoring work.
Filters
Usage volume (ROI)
Priority
Model Ranking (14 models)
Est. Monthly Cost is a total spend estimate at 10,000 requests/mo × 2,000 in + 500 out tokens — not a per-1M-token rate. Per-token list prices are in the In / 1M and Out / 1M columns.
| Task rank# | Select | Model | License | Badges | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 ProDeepSeek | 80 | 82 | 4478 tok/s | 1.0M100 | $0.43 | $0.87 | $13.05 | DeepSeek | ||
| 2 | DeepSeek V4 FlashDeepSeek | 75 | 60 | 82168 tok/s | 1.0M100 | $0.14 | $0.28 | $4.20 | DeepSeek | ||
| 3 | Qwen 3.7 MaxQwen | 74 | 79 | 4478 tok/s | 1M100 | $1.48 | $4.42 | $51.63 | Qwen-Research | ||
| 4 | GPT-OSS 120BOpenAI | 74 | 59 | 77148 tok/s | 131K100 | $0.037 | $0.17 | $1.59 | Apache-2.0 | ||
| 5 | Claude Haiku 4.5Anthropic | 72 | 70 | 59108 tok/s | 200K100 | $1.00 | $5.00 | $45.00 | Proprietary | ||
| 6 | Gemma 4 31BGoogle | 72 | 56 | 77148 tok/s | 262K100 | $0.10 | $0.34 | $3.70 | Gemma | ||
| 7 | MiMo-V2.5Xiaomi | 72 | 56 | 77148 tok/s | 1.1M100 | $0.14 | $0.28 | $4.20 | Open-Weights-Other | ||
| 8 | Qwen 3.7 PlusQwen | 71 | 56 | 77148 tok/s | 1M100 | $0.32 | $1.28 | $12.80 | Qwen-Research | ||
| 9 | Claude Sonnet 5Anthropic | 70 | 80 | 4478 tok/s | 1M100 | $2.00 | $10.00 | $90.00 | Proprietary | ||
| 10 | Qwen 3.7 FlashQwen | 68 | 41 | 100210 tok/s | 1M100 | $0.030 | $0.13 | $1.25 | Qwen-Research | ||
| 11 | Nemotron 3 Ultra (free)NVIDIA | 67 | 41 | 96195 tok/s | 1M100 | Free | Free | $0.00 | Llama-Community | ||
| 12 | Claude Opus 5Anthropic | 58 | 86 | 3360 tok/s | 1M100 | $5.00 | $25.00 | $225.00 | Proprietary | ||
| 13 | DevTools Code Quality (1.5B)DevTools (self-hosted, Qwen2.5-Coder base) | 43 | 32 | 028 tok/s | 4K6 | Free | Free | $0.00 | Apache-2.0 | ||
| 14 | DevTools Code Light (0.5B)DevTools (self-hosted, Qwen2.5-Coder base) | 43 | 24 | 2855 tok/s | 4K6 | Free | Free | $0.00 | Apache-2.0 |
Pricing, context windows and model names come from the same catalog as the LLM token counter. Benchmark scores are hand-reviewed reference values based on public leaderboards (LM Arena, SWE-bench), not a live leaderboard feed — last updated 2026-08-20.
Related tools
Agent Tool Flowchart Builder
Design agentic tool-call DAGs with decision branches and bounded loops. Validates reachability, cycle handling, and emits clean Mermaid flowcharts — fully in your browser.
AI Agent Memory & Buffer Simulator
Compare sliding window, summary buffer, and vector memory strategies
AI API Client SDK Generator
Generate typed API client SDK files from OpenAPI specifications or cURL requests
AI Code Explainer
Explain code snippets with grounded cloud AI
AI Code Security Auditor
Scan source code for high-confidence OWASP Top 10 vulnerabilities
AI Code Translator
Translate code into idiomatic implementations in another language