LLM Benchmarks
Compare LLM coding, reasoning, speed and cost benchmarks side-by-side and get a task-specific model recommendation
Task
Weighted toward LiveCodeBench, SWE-bench Lite and Aider — for agentic coding and refactoring work.
Filters
Usage volume (ROI)
Priority
Model Ranking (14 models)
Est. Monthly Cost is a total spend estimate at 10,000 requests/mo × 2,000 in + 500 out tokens — not a per-1M-token rate. Per-token list prices are in the In / 1M and Out / 1M columns.
| Task rank# | Select | Model | License | Badges | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 ProDeepSeek | 80 | 82 | 4478 tok/s | 1.0M100 | $0.43 | $0.87 | $13.05 | DeepSeek | ||
| 2 | DeepSeek V4 FlashDeepSeek | 75 | 60 | 82168 tok/s | 1.0M100 | $0.14 | $0.28 | $4.20 | DeepSeek | ||
| 3 | Qwen 3.7 MaxQwen | 74 | 79 | 4478 tok/s | 1M100 | $1.48 | $4.42 | $51.63 | Qwen-Research | ||
| 4 | GPT-OSS 120BOpenAI | 74 | 59 | 77148 tok/s | 131K100 | $0.037 | $0.17 | $1.59 | Apache-2.0 | ||
| 5 | Claude Haiku 4.5Anthropic | 72 | 70 | 59108 tok/s | 200K100 | $1.00 | $5.00 | $45.00 | Proprietary | ||
| 6 | Gemma 4 31BGoogle | 72 | 56 | 77148 tok/s | 262K100 | $0.10 | $0.34 | $3.70 | Gemma | ||
| 7 | MiMo-V2.5Xiaomi | 72 | 56 | 77148 tok/s | 1.1M100 | $0.14 | $0.28 | $4.20 | Open-Weights-Other | ||
| 8 | Qwen 3.7 PlusQwen | 71 | 56 | 77148 tok/s | 1M100 | $0.32 | $1.28 | $12.80 | Qwen-Research | ||
| 9 | Claude Sonnet 5Anthropic | 70 | 80 | 4478 tok/s | 1M100 | $2.00 | $10.00 | $90.00 | Proprietary | ||
| 10 | Qwen 3.7 FlashQwen | 68 | 41 | 100210 tok/s | 1M100 | $0.030 | $0.13 | $1.25 | Qwen-Research | ||
| 11 | Nemotron 3 Ultra (free)NVIDIA | 67 | 41 | 96195 tok/s | 1M100 | Free | Free | $0.00 | Llama-Community | ||
| 12 | Claude Opus 5Anthropic | 58 | 86 | 3360 tok/s | 1M100 | $5.00 | $25.00 | $225.00 | Proprietary | ||
| 13 | DevTools Code Quality (1.5B)DevTools (self-hosted, Qwen2.5-Coder base) | 43 | 32 | 028 tok/s | 4K6 | Free | Free | $0.00 | Apache-2.0 | ||
| 14 | DevTools Code Light (0.5B)DevTools (self-hosted, Qwen2.5-Coder base) | 43 | 24 | 2855 tok/s | 4K6 | Free | Free | $0.00 | Apache-2.0 |
Pricing, context windows and model names come from the same catalog as the LLM token counter. Benchmark scores are hand-reviewed reference values based on public leaderboards (LM Arena, SWE-bench), not a live leaderboard feed — last updated 2026-08-20.
Examples
Find the best model for agentic coding
Open the Coding tab and read the Best Value badge — the ranking leans on LiveCodeBench, SWE-bench Lite and Aider scores at your chosen request volume.
Budget a high-volume chat feature
Switch to the Fast Chat tab, enter your expected monthly requests and average tokens per request, and slide the priority toward Savings.
Try a fully offline browser model
Open the Edge/Browser tab, check a WebLLM model's download size and VRAM requirement, then use its "Try in browser" link.
About this tool
Picking a model isn't just about a single leaderboard number — the best choice for an agentic coding task is rarely the best choice for a cheap high-volume chatbot or a fully offline browser tool. This matrix scores every model against the benchmarks that actually matter for five task types (coding, documentation/RAG, reasoning, edge/browser and fast chat), weighted differently per task, and rolls them into one comparable Task Score.
Pick a task tab to filter the table to relevant models and see the weighting behind the score. Enter your own monthly request volume and average input/output tokens to see an estimated monthly cost per model, and slide the quality-vs-savings priority to re-rank instantly. Select up to 4 models to compare on a radar chart across quality, cost, speed, context and (for edge models) memory footprint.
Every model card is source-linked (LM Arena, SWE-bench) and dated. Benchmark scores are hand-reviewed reference values, not a live leaderboard feed — this stays a static, client-side snapshot with no network calls in this phase.
How to use
Pick a task
Choose Coding, Documentation & RAG, Reasoning, Edge/Browser or Fast Chat — the weighting behind the score changes per task.
Enter your traffic
Type your expected monthly requests and average input/output tokens per request to see an estimated monthly cost per model.
Slide the priority
Move the quality-vs-savings slider to instantly re-rank models toward maximum quality or maximum cost savings.
Compare on a radar chart
Check up to 4 models and open Compare to see quality, cost, speed, context and memory plotted side-by-side.
Use cases
Model selection for a new feature
Compare quality, speed and cost together instead of chasing a single leaderboard number that doesn't match your task.
Budgeting a chatbot or agent
Enter real traffic assumptions to see estimated monthly cost per model before committing to a provider.
Evaluating on-device/offline AI
Filter to Edge/Browser to compare WebLLM models on download size, VRAM and quality trade-offs.
Common mistakes
Mistake:Comparing raw benchmark numbers across different task types
Fix:Use the task-specific Task Score instead — it's normalized and weighted for what actually matters for that kind of work.
Mistake:Treating this as a live, auto-updating leaderboard
Fix:Scores are a dated, hand-reviewed snapshot (see the last-updated note) — re-check official sources for the latest numbers before a high-stakes decision.
Frequently asked questions
References & standards
Related tools
OCR — Image to Text Extractor
Extract text from screenshots, scans and photos with a self-hosted OCR engine — image enhanced automatically, nothing uploaded.
Video to GIF Maker
Cut a clip from any video and convert it to a high-quality GIF — palette-optimized colors, size and FPS control, fully local.
Currency Converter
Convert between 30 currencies using an offline rate snapshot with manual rate override — no API calls, works without a network.
Text to Speech (TTS)
Read any text aloud with your browser's built-in voices — control voice, rate, pitch and volume, all processed on your device.
Voice Dictation
Dictate with your voice and watch the text appear live — 12 languages, browser speech engine, transcript stays on your device.
Ambient Noise Generator
Blend white, pink and brown noise live in your browser — focus masking, tinnitus relief and sleep, with a built-in sleep timer.