EXPERT INSIGHTS

LLM leaderboard

Use model benchmarks and specifications to build a shortlist. Then test the candidates on the documents, questions, and tool calls your application will handle.


A public score is a starting point. It does not establish how a model will perform in your workflow.

MMLU Benchmarks

MMLU evaluates knowledge and problem-solving across 57 academic and professional tasks. It is not a benchmark limited to grade-school mathematics.

HumanEval+

Extended version of HumanEval with more complex programming challenges across multiple languages to test code quality.

GPQA Evaluation

Graduate-level expert knowledge evaluation designed to test advanced reasoning in specialized domains.

MT-Bench Analysis

Multi-turn benchmarking that evaluates conversation abilities, reasoning, and instruction following across complex dialogues.

SWE Benchmarks

Software engineering tests including code generation, debugging, and algorithm design to measure programming capabilities.

GSM8K Reasoning

Grade school math word problems requiring multi-step reasoning to evaluate logical thinking and problem-solving capabilities.

Compare the measurements that matter


Choose the models and measurements relevant to your task. Before relying on a score or specification, check the original source, model version, and measurement conditions.


Treat an unavailable value as missing information, not as a zero. Do not infer that a model supports a capability merely because another model from the same provider does.

Provider

Model Name

MMLU Score

Parameters

Context Size

Price

AI21

Jamba Large 1.7

N/A

N/A

256,000

3.5

AI21

Jamba Mini 2

N/A

N/A

256,000

0.25

Amazon

Nova 2 Lite

80.9%

N/A

1,000,000

0.85

Amazon

Nova 2 Omni

N/A

N/A

1,000,000

0.85

Amazon

Nova 2 Pro

N/A

N/A

1,000,000

3.438

Amazon

Nova 2 Sonic

N/A

N/A

1,000,000

0.935

Amazon

Nova Premier

N/A

N/A

1,000,000

5

Amazon

Nova Pro

69.1%

N/A

300,000

1.4

Anthropic

Claude 3.5 Haiku

63.4%

N/A

200,000

1.6

Anthropic

Claude 3.7 Sonnet

80.3%

N/A

200,000

7

Test your shortlist on real work


Choose representative examples and define an acceptable result before running the comparison. Keep the inputs and evaluation rules consistent across models.


Track incorrect answers, missing evidence, invalid tool calls, elapsed time, and the cost of obtaining an accepted result. Include cases where the system should ask for clarification or stop.


Need to evaluate a complete workflow?


Bring a sample input, the expected output, and the systems involved. We can discuss how to test the model and the surrounding workflow together.


Discuss your AI workflow

MMLU Scores

Compare results for the benchmark version and evaluation setting recorded with each entry.

MMLU measures general knowledge and reasoning

MMLU measures general knowledge and reasoning

Claude 4.7 Opus

MMLU Score

92.8%

GPT-5.5

MMLU Score

92.4%

Gemini 3 Pro

MMLU Score

89.8%

Gemini 3.1 Flash Lite

MMLU Score

89.2%

Claude 4.5 Opus

MMLU Score

88.9%

MMLU measures general knowledge and reasoning

MMLU measures general knowledge and reasoning

Claude 4.7 Opus

Claude 4.7 Opus

Claude 4.7 Opus

GPT-5.5

GPT-5.5

GPT-5.5

Claude 4 Opus

Claude 4 Opus

Claude 4 Opus

Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite

Claude 4.5 Opus

Claude 4.5 Opus

Claude 4.5 Opus

80%

90%

100%

Reported Token Throughput

Compare measurements only after checking the provider, test conditions, and measurement date.

Tokens processed per second - higher is better

Tokens processed per second - higher is better

Mercury 2

Throughput

870.9

tokens/s

Inference Speed

0

ms/tokens

Latency

3.67

ms

Provider

Inception

Granite 4.0 Small

Throughput

448.8

tokens/s

Inference Speed

0

ms/tokens

Latency

0.55

ms

Provider

IBM

Nemotron 3 Super 120B

Throughput

390.7

tokens/s

Inference Speed

0

ms/tokens

Latency

0.7

ms

Provider

NVIDIA

GPT-OSS 120B

Throughput

319.541

tokens/s

Inference Speed

0

ms/tokens

Latency

0.47

ms

Provider

OpenAI

GPT-OSS 20B

Throughput

297.704

tokens/s

Inference Speed

0

ms/tokens

Latency

0.503

ms

Provider

OpenAI

Tokens processed per second - higher is better

Mercury 2

Mercury 2

Granite 4.0 Small

Granite 4.0 Small

Nemotron 3 Super 120B

Nemotron 3 Super 120B

GPT-OSS 120B

GPT-OSS 120B

GPT-OSS 20B

GPT-OSS 20B

0.0

500

1.000

Published Model Pricing

Check input and output prices separately, including the billing unit and any conditions attached to the rate.

Price per million tokens - lower is better

Gemma 3n E4B

Gemma 3n E4B

Llama 3.2 1B Instruct

Llama 3.2 1B Instruct

Command R7B

Command R7B

Granite 4.0 Small

Granite 4.0 Small

Qwen 3.5 9B

Qwen 3.5 9B

$0.000

$0.040

$0.080

Price per million tokens - lower is better

Price per million tokens - lower is better

Gemma 3n E4B

Input Price

0.03

$/M

Output Price

0.06

$/M

Effective Price

0.037

$/M

Provider

Google

Llama 3.2 1B Instruct

Input Price

0.053

$/M

Output Price

0.055

$/M

Effective Price

0.053

$/M

Provider

Meta

Command R7B

Input Price

0.0375

$/M

Output Price

0.15

$/M

Effective Price

0.066

$/M

Provider

Cohere

Granite 4.0 Small

Input Price

0.05

$/M

Output Price

0.15

$/M

Effective Price

0.075

$/M

Provider

IBM

Qwen 3.5 9B

Input Price

0.05

$/M

Output Price

0.15

$/M

Effective Price

0.075

$/M

Provider

Qwen

Published Context Limits

Check the exact model and endpoint before using a context limit in your application design.

While tokenization varies between models, on average, 1 token ≈ 3.5 characters in English

Grok 4 Fast

Grok 4 Fast

Grok 4 Heavy

Grok 4 Heavy

Grok 4.1

Grok 4.1

Grok 4.20

Grok 4.20

GPT 4.1

GPT 4.1

0

1M

2M

While tokenization varies between models, on average, 1 token ≈ 3.5 characters in English

While tokenization varies between models, on average, 1 token ≈ 3.5 characters in English

Grok 4 Fast

Tokens

2,000,000

Grok 4 Heavy

Tokens

2,000,000

Grok 4.1

Tokens

2,000,000

Grok 4.20

Tokens

2,000,000

GPT 4.1

Tokens

1,280,000

What different token speeds look like

What different token speeds look like

What different token speeds look like

What different token speeds look like

This animation illustrates example speeds. It is not a live measurement of a model or provider.

1200

t/s

The quick brown fox jumps over the lazy dog. Meanwhile, a clever rabbit watches from nearby bushes, intrigued by the scene unfolding before its eyes. The fox continues its playful pursuit, demonstrating remarkable agility and grace in motion. As the sun sets on the horizon, the forest comes alive with the sounds of nature, creating a symphony of rustling leaves and gentle breezes. The fox pauses, alert to these changes, its ears perked up to catch every subtle noise in the surroundings.

200

t/s

The quick brown fox jumps over the lazy dog. Meanwhile, a clever rabbit watches from nearby bushes, intrigued by the scene unfolding before its eyes. The fox continues its playful pursuit, demonstrating remarkable agility and grace in motion. As the sun sets on the horizon, the forest comes alive with the sounds of nature, creating a symphony of rustling leaves and gentle breezes. The fox pauses, alert to these changes, its ears perked up to catch every subtle noise in the surroundings.

40

t/s

The quick brown fox jumps over the lazy dog. Meanwhile, a clever rabbit watches from nearby bushes, intrigued by the scene unfolding before its eyes. The fox continues its playful pursuit, demonstrating remarkable agility and grace in motion. As the sun sets on the horizon, the forest comes alive with the sounds of nature, creating a symphony of rustling leaves and gentle breezes. The fox pauses, alert to these changes, its ears perked up to catch every subtle noise in the surroundings.

Values reset every 5 seconds to demonstrate different speeds

Compare LLM Models

Compare LLM Models

Compare LLM Models

Compare LLM Models

Compare any two LLM models side by side across different metrics, including MMLU, GPQA, HumanEval, DROP, Context Size, Parameters, Input Price, Output Price, Inference Speed, Throughput, and Latency.

Metric

Provider

MMLU Score

GPQA Score

HumanEval Score

Context Size

Parameters

Input Price

Throughput

Latency

Claude 3.5 Haiku

Anthropic

63.4%

40.8%

75.6%

200,000

N/A

0.8

49.093

0.689

Claude 3.7 Sonnet

Anthropic

80.3%

65.6%

92.1%

200,000

N/A

3

N/A

N/A

Get started

Let’s Build AI Agents, Together

Book a demo to see how AI agents can help your team process unstructured documents and perform complex analysis faster and more accurately.

StackAI cube logo mark
Dark rounded AI processor chip illustration

Get started

Let’s Build AI Agents, Together

Book a demo to see how AI agents can help your team process unstructured documents and perform complex analysis faster and more accurately.

StackAI cube logo mark
Dark rounded AI processor chip illustration